English
Related papers

Related papers: A High-Performance Design, Implementation, Deploym…

200 papers

We introduce a high-performance cost-effective network topology called Slim Fly that approaches the theoretically optimal network diameter. Slim Fly is based on graphs that approximate the solution to the degree-diameter problem. We analyze…

Networking and Internet Architecture · Computer Science 2020-07-01 Maciej Besta , Torsten Hoefler

We introduce FatPaths: a simple, generic, and robust routing architecture that enables state-of-the-art low-diameter topologies such as Slim Fly to achieve unprecedented performance. FatPaths targets Ethernet stacks in both HPC…

Networking and Internet Architecture · Computer Science 2020-11-12 Maciej Besta , Marcel Schneider , Karolina Cynk , Marek Konieczny , Erik Henriksson , Salvatore Di Girolamo , Ankit Singla , Torsten Hoefler

Emerging chips with hundreds and thousands of cores require networks with unprecedented energy/area efficiency and scalability. To address this, we propose Slim NoC (SN): a new on-chip network design that delivers significant improvements…

Hardware Architecture · Computer Science 2020-10-22 Maciej Besta , Syed Minhaj Hassan , Sudhakar Yalamanchili , Rachata Ausavarungnirun , Onur Mutlu , Torsten Hoefler

The recent line of research into topology design focuses on lowering network diameter. Many low-diameter topologies such as Slim Fly or Jellyfish that substantially reduce cost, power consumption, and latency have been proposed. A key…

Networking and Internet Architecture · Computer Science 2020-11-02 Maciej Besta , Jens Domke , Marcel Schneider , Marek Konieczny , Salvatore Di Girolamo , Timo Schneider , Ankit Singla , Torsten Hoefler

Extreme-scale data centers are the backbone of next-generation computing, enabling breakthroughs in science, artificial intelligence, and global innovation through unprecedented processing power and scalability. This work examines…

Networking and Internet Architecture · Computer Science 2026-05-27 Alejandro Cano , Cristina Brinza , Cristóbal Camarero , Carmen Martínez , Ramón Beivide

Low-diameter topologies such as Dragonfly and Slim Fly are increasingly adopted in HPC and datacenter networks, yet existing load balancing techniques either rely on proprietary in-network mechanisms or fail to utilize the full path…

Networking and Internet Architecture · Computer Science 2026-02-24 Tommaso Bonato , Ales Kubicek , Abdul Kabbani , Ahmad Ghalayini , Maciej Besta , Torsten Hoefler

Mobile devices are indispensable sources of big data. Federated learning (FL) has a great potential in exploiting these private data by exchanging locally trained models instead of their raw data. However, mobile devices are often energy…

Machine Learning · Computer Science 2021-12-08 Hankyul Baek , Won Joon Yun , Soyi Jung , Jihong Park , Mingyue Ji , Joongheon Kim , Mehdi Bennis

Multi-plane architectures have become increasingly prevalent in the Fat-Tree networks of AI data centers. By leveraging multiple ports on a single network interface card (NIC) or multiple NICs within a scale-up domain, each port or NIC is…

Networking and Internet Architecture · Computer Science 2026-04-28 Ziyu Wang , Fei Lei , Dezun Dong

Industry experience indicates that the ability to incrementally expand data centers is essential. However, existing high-bandwidth network designs have rigid structure that interferes with incremental expansion. We present Jellyfish, a…

Networking and Internet Architecture · Computer Science 2012-04-24 Ankit Singla , Chi-Yao Hong , Lucian Popa , P. Brighten Godfrey

Existing high-performance computing (HPC) interconnection architectures are based on high-radix switches, which limits the injection/local performance and introduces latency/energy/cost overhead. The new wafer-scale packaging and high-speed…

Hardware Architecture · Computer Science 2024-08-27 Yinxiao Feng , Kaisheng Ma

Training large and highly accurate deep learning (DL) models is computationally costly. This cost is in great part due to the excessive number of trained parameters, which are well-known to be redundant and compressible for the execution…

Machine Learning · Computer Science 2019-04-11 Mojan Javaheripi , Bita Darvish Rouhani , Farinaz Koushanfar

Federated learning (FL) is a key enabler for efficient communication and computing, leveraging devices' distributed computing capabilities. However, applying FL in practice is challenging due to the local devices' heterogeneous energy,…

Machine Learning · Computer Science 2022-12-23 Won Joon Yun , Yunseok Kwak , Hankyul Baek , Soyi Jung , Mingyue Ji , Mehdi Bennis , Jihong Park , Joongheon Kim

Convolutional Neural Networks (CNNs) have demonstrated their effectiveness in numerous vision tasks. However, their high processing requirements necessitate efficient hardware acceleration to meet the application's performance targets. In…

Hardware Architecture · Computer Science 2024-03-29 Petros Toupas , Zhewen Yu , Christos-Savvas Bouganis , Dimitrios Tzovaras

We design and deploy in production the first flat datacenter networks. Our design, called RNG, is based on quasi-random graphs. While the cost and fault-tolerance benefits of such topologies have been long known, their practical realization…

Increasingly large AI workloads are calling for hyper-scale infrastructure; however, traditional interconnection network architecture is neither scalable nor cost-effective enough. Tree-based topologies such as the \textit{Rail-optimized}…

Hardware Architecture · Computer Science 2025-07-28 Yinxiao Feng , Tiancheng Chen , Yuchen Wei , Siyuan Shen , Shiju Wang , Wei Li , Kaisheng Ma , Torsten Hoefler

Recent work has shown that expander-based data center topologies are robust and can yield superior performance over Clos topologies. However, to achieve these benefits, previous proposals use routing and transport schemes that impede quick…

Networking and Internet Architecture · Computer Science 2018-11-02 Vipul Harsh , Sangeetha Abdu Jyothi , Inderdeep Singh , P. Brighten Godfrey

Cross-silo Federated Learning (FL) enables multiple institutions to collaboratively train machine learning models while preserving data privacy. In such settings, clients repeatedly exchange model weights with a central server, making the…

Networking and Internet Architecture · Computer Science 2025-09-05 Osama Abu Hamdan , Hao Che , Engin Arslan , Md Arifuzzaman

The diversity of communication paths in a network, especially non-minimal paths, is a key enabler of performance at extreme scales. We present EvalNet, a toolchain for scalable generation and analysis of over 25 important network…

Current dynamic networks and dynamic pruning methods have shown their promising capability in reducing theoretical computation complexity. However, dynamic sparse patterns on convolutional filters fail to achieve actual acceleration in…

Computer Vision and Pattern Recognition · Computer Science 2021-03-25 Changlin Li , Guangrun Wang , Bing Wang , Xiaodan Liang , Zhihui Li , Xiaojun Chang

Despite the latest prevailing success of deep neural networks (DNNs), several concerns have been raised against their usage, including the lack of intepretability the gap between DNNs and other well-established machine learning models, and…

Machine Learning · Computer Science 2021-01-01 Jianghao Shen , Sicheng Wang , Zhangyang Wang
‹ Prev 1 2 3 10 Next ›