English
Related papers

Related papers: RailX: A Flexible, Scalable, and Low-Cost Network …

200 papers

This paper presents a low-cost network architecture for training large language models (LLMs) at hyperscale. We study the optimal parallelization strategy of LLMs and propose a novel datacenter network design tailored to LLM's unique…

Networking and Internet Architecture · Computer Science 2024-09-17 Weiyang Wang , Manya Ghobadi , Kayvon Shakeri , Ying Zhang , Naader Hasani

As high-performance computing systems scale in size and complexity, efficient resource management is essential to minimize communication overhead. The HyperX is a richly connected, low-diameter network that offers a scalable and…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-28 Alejandro Cano , Cristóbal Camarero , Carmen Martínez , Ramón Beivide

Multi-plane architectures have become increasingly prevalent in the Fat-Tree networks of AI data centers. By leveraging multiple ports on a single network interface card (NIC) or multiple NICs within a scale-up domain, each port or NIC is…

Networking and Internet Architecture · Computer Science 2026-04-28 Ziyu Wang , Fei Lei , Dezun Dong

In modern network design, "efficiency" is often conflated with raw performance metrics like latency or aggregate throughput. This paper proposes a resource-centric definition of efficiency, isolating the hardware cost required to maintain a…

Networking and Internet Architecture · Computer Science 2026-01-28 Jia Xu Wei , Wei Wei

We propose TopoOpt, a novel direct-connect fabric for deep neural network (DNN) training workloads. TopoOpt co-optimizes the distributed training process across three dimensions: computation, communication, and network topology. We…

Networking and Internet Architecture · Computer Science 2022-10-03 Weiyang Wang , Moein Khazraee , Zhizhen Zhong , Manya Ghobadi , Zhihao Jia , Dheevatsa Mudigere , Ying Zhang , Anthony Kewitsch

Extreme-scale data centers are the backbone of next-generation computing, enabling breakthroughs in science, artificial intelligence, and global innovation through unprecedented processing power and scalability. This work examines…

Networking and Internet Architecture · Computer Science 2026-05-27 Alejandro Cano , Cristina Brinza , Cristóbal Camarero , Carmen Martínez , Ramón Beivide

Interconnection networks are key actors that condition the performance of current large datacenter and supercomputer systems. Both topology and routing are critical aspects that must be carefully considered for a competitive system network…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-04-09 Cristóbal Camarero , Alejandro Cano , Carmen Martínez , Ramón Beivide

Rail-optimized network fabrics have become the de facto datacenter scale-out fabric for large-scale ML training. However, the use of high-radix electrical switches to provide all-to-all connectivity in rails imposes massive power, cost, and…

Networking and Internet Architecture · Computer Science 2025-07-14 Eric Ding , Chuhan Ouyang , Rachee Singh

We introduce LilNetX, an end-to-end trainable technique for neural networks that enables learning models with specified accuracy-rate-computation trade-off. Prior works approach these problems one at a time and often require post-processing…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Sharath Girish , Kamal Gupta , Saurabh Singh , Abhinav Shrivastava

Rail-optimized network fabrics have become the de facto datacenter scale-out fabric for large-scale ML training. However, the use of high-radix electrical switches to provide all-to-all connectivity in rails imposes massive power and cost.…

Networking and Internet Architecture · Computer Science 2026-03-18 Eric Ding , Barry Lyu , Bhaskar Kataria , Rachee Singh

The interconnection network comprises a significant portion of the cost of large parallel computers, both in economic terms and power consumption. Several previous proposals exploit large-radix routers to build scalable low-distance…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-12-24 Cristóbal Camarero , Carmen Martínez , Enrique Vallejo , Ramón Beivide

Existing high-performance computing (HPC) interconnection architectures are based on high-radix switches, which limits the injection/local performance and introduces latency/energy/cost overhead. The new wafer-scale packaging and high-speed…

Hardware Architecture · Computer Science 2024-08-27 Yinxiao Feng , Kaisheng Ma

Distributed deep learning (DDL) systems strongly depend on network performance. Current electronic packet switched (EPS) network architectures and technologies suffer from variable diameter topologies, low-bisection bandwidth and…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-02-27 Alessandro Ottino , Joshua Benjamin , Georgios Zervas

Hierarchical ring networks, which hierarchically connect multiple levels of rings, have been proposed in the past to improve the scalability of ring interconnects, but past hierarchical ring designs sacrifice some of the key benefits of…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-02-22 Rachata Ausavarungnirun , Chris Fallin , Xiangyao Yu , Kevin Kai-Wei Chang , Greg Nazario , Reetuparna Das , Gabriel H. Loh , Onur Mutlu

Numerous microarchitectural optimizations unlocked tremendous processing power for deep neural networks that in turn fueled the AI revolution. With the exhaustion of such optimizations, the growth of modern AI is now gated by the…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-10-24 Torsten Hoefler , Tommaso Bonato , Daniele De Sensi , Salvatore Di Girolamo , Shigang Li , Marco Heddes , Jon Belk , Deepak Goel , Miguel Castro , Steve Scott

We introduce a high-performance cost-effective network topology called Slim Fly that approaches the theoretically optimal network diameter. Slim Fly is based on graphs that approximate the solution to the degree-diameter problem. We analyze…

Networking and Internet Architecture · Computer Science 2020-07-01 Maciej Besta , Torsten Hoefler

Novel low-diameter network topologies such as Slim Fly (SF) offer significant cost and power advantages over the established Fat Tree, Clos, or Dragonfly. To spearhead the adoption of low-diameter networks, we design, implement, deploy, and…

Off-chain transaction channels represent one of the leading techniques to scale the transaction throughput in cryptocurrencies such as Bitcoin. They allow multiple agents to route payments through one another. So far, the topology and…

Computer Science and Game Theory · Computer Science 2020-07-28 Yotam Sali , Aviv Zohar

We present some novel, straightforward methods for training the connection graph of a randomly initialized neural network without training the weights. These methods do not use hyperparameters defining cutoff thresholds and therefore remove…

Machine Learning · Computer Science 2020-11-18 Cristian Ivan , Razvan Florian

As distributed model training scales to span hundreds of thousands of GPUs, scale-out networks face unprecedented performance and efficiency demands. NVIDIA Spectrum-X Ethernet has been designed from the ground up to achieve predictable and…

‹ Prev 1 2 3 10 Next ›