中文
相关论文

相关论文: Canary: Congestion-Aware In-Network Allreduce Usin…

200 篇论文

Ring-based collective operations are widely used in distributed AI training due to their efficient bandwidth utilization. While ring communication excels at pipelining, its performance is heavily dependent on having synchronized step-wise…

网络与互联网体系结构 · 计算机科学 2026-04-21 Yuze Jin , Xin Zhe Khooi , Ruyi Yao , Mun Choon Chan

The centralized architecture in software-defined network (SDN) provides a global view of the underlying network, paving the way for enormous research in the area of SDN traffic engineering (SDN TE). This research focuses on the load…

网络与互联网体系结构 · 计算机科学 2018-12-07 Sminesh C. N. , Grace Mary Kanaga E. , Ranjitha K

As communication protocols evolve, datacenter network utilization increases. As a result, congestion is more frequent, causing higher latency and packet loss. Combined with the increasing complexity of workloads, manual design of congestion…

网络与互联网体系结构 · 计算机科学 2024-06-04 Benjamin Fuhrer , Yuval Shpigelman , Chen Tessler , Shie Mannor , Gal Chechik , Eitan Zahavi , Gal Dalal

The evolution of wireless networks and radio access technologies (RATs) has transformed communication from user-driven traffic into a dynamic ecosystem of autonomous systems, including IoT devices, edge nodes, autonomous vehicles, AR/XR…

网络与互联网体系结构 · 计算机科学 2025-11-04 Pavan K. Mangipudi , Sharon Boamah , Lorenz Carvajal , Janise Mcnair

In this work, we provide the design and implementation of a switch-assisted congestion control algorithm for data center networks (DCNs). In particular, we provide a prototype of the switch-driven congestion control algorithm and deploy it…

网络与互联网体系结构 · 计算机科学 2021-06-29 Ahmed M. Abdelmoniem , Brahim Bensaou

Emerging artificial intelligence (AI) and machine learning (ML) workloads present new challenges of managing the collective communication used in distributed training across hundreds or even thousands of GPUs. This paper presents STrack, a…

网络与互联网体系结构 · 计算机科学 2024-07-25 Yanfang Le , Rong Pan , Peter Newman , Jeremias Blendin , Abdul Kabbani , Vipin Jain , Raghava Sivaramu , Francis Matus

The method of choice for parameter aggregation in Deep Neural Network (DNN) training, a network-intensive task, is shifting from the Parameter Server model to decentralized aggregation schemes (AllReduce) inspired by theoretical guarantees…

网络与互联网体系结构 · 计算机科学 2020-04-30 Sayed Hadi Hashemi , Sangeetha Abdu Jyothi , Brighten Godfrey , Roy Campbell

We present a unified analytical framework within which power control, rate allocation, routing, and congestion control for wireless networks can be optimized in a coherent and integrated manner. We consider a multi-commodity flow model with…

网络与互联网体系结构 · 计算机科学 2007-09-18 Yufang Xi , Edmund M. Yeh

Edge machine learning can deliver low-latency and private artificial intelligent (AI) services for mobile devices by leveraging computation and storage resources at the network edge. This paper presents an energy-efficient edge processing…

信息论 · 计算机科学 2020-03-03 Kai Yang , Yuanming Shi , Wei Yu , Zhi Ding

We focus on the commonly used synchronous Gradient Descent paradigm for large-scale distributed learning, for which there has been a growing interest to develop efficient and robust gradient aggregation strategies that overcome two key…

The interconnection network is a key element in High-Performance Computing (HPC) and Datacenter (DC) systems whose performance depends on several design parameters, such as the topology, the switch architecture, and the routing algorithm.…

网络与互联网体系结构 · 计算机科学 2025-02-04 Jose Rocher-Gonzalez , Jesus Escudero-Sahuquillo , Pedro J. Garcia , Francisco J. Quiles , Gaspar Mora

This paper presents an asynchronous distributed algorithm to manage multiple trees for peer-to-peer streaming in a flow level model. It is assumed that videos are cut into substreams, with or without source coding, to be distributed to all…

数据结构与算法 · 计算机科学 2013-08-12 Ji Zhu , Bruce Hajek

For highly distributed environments such as edge computing, collaborative learning approaches eschew the dependence on a global, shared model, in favor of models tailored for each location. Creating tailored models for individual learning…

机器学习 · 计算机科学 2021-08-31 Harshit Daga , Yiwen Chen , Aastha Agrawal , Ada Gavrilovska

In recent years, the issue of energy consumption in high performance computing (HPC) systems has attracted a great deal of attention. In response to this, many energy-aware algorithms have been developed in different layers of HPC systems,…

分布式、并行与集群计算 · 计算机科学 2014-05-13 Nikzad Babaii Rizvandi

We consider a network of smart sensors for an edge computing application that sample a time-varying signal and send updates to a base station for remote global monitoring. Sensors are equipped with sensing and compute, and can either send…

分布式、并行与集群计算 · 计算机科学 2025-02-11 Luca Ballotta , Giovanni Peserico , Francesco Zanini , Paolo Dini

Wirelessly connected vehicles that exchange information about traffic conditions can reduce delays caused by congestion. At a 2-to-1 lane reduction, the improvement in flow past a bottleneck due to traffic with a random mixture of 40%…

物理与社会 · 物理学 2016-03-23 L. C. Davis

Being able to identify service slowdowns is crucial to many operational problems. We study how to use observational congestion data to learn service slowdown in a multi-server system that uses adaptive congestion control mechanisms. We show…

物理与社会 · 物理学 2025-03-18 Xu Kuang , Gal Mendelson

Large inter-GPU all-reduce operations, prevalent throughout deep learning, are bottlenecked by communication costs. Emerging heterogeneous architectures are comprised of complex nodes, often containing $4$ GPUs and dozens to hundreds of CPU…

分布式、并行与集群计算 · 计算机科学 2026-02-26 Michael Adams , Amanda Bienz

Emerging edge computing paradigms enable heterogeneous devices to collaborate on complex computation applications. However, for congestible links and computing units, delay-optimal forwarding and offloading for service chain tasks (e.g.,…

网络与互联网体系结构 · 计算机科学 2024-03-26 Jinkun Zhang , Yuezhou Liu , Edmund Yeh

Communication scheduling has been shown to be effective in accelerating distributed training, which enables all-reduce communications to be overlapped with backpropagation computations. This has been commonly adopted in popular distributed…

机器学习 · 计算机科学 2023-06-16 Lin Zhang , Shaohuai Shi , Xiaowen Chu , Wei Wang , Bo Li , Chengjian Liu