English
Related papers

Related papers: Resilient AI Supercomputer Networking using MRC an…

200 papers

Datacenter network design plays a critical role in AI training by supporting scaling to thousands of accelerators. An open problem, designing a near-optimal throughput oriented network-topology, routing, and collectives-has not been…

Networking and Internet Architecture · Computer Science 2026-05-28 Conor James Green , Mithuna Thottethodi

In this paper, we develop RCC, the first unified and comprehensive RDMA-enabled distributed transaction processing framework supporting six serializable concurrency control protocols: not only the classical protocols NOWAIT, WAITDIE, and…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-01-21 Chao Wang , Kezhao Huang , Xuehai Qian

Recent advances in multimodal agents have improved computer-use interaction and tool-usage, yet most existing systems remain reactive, optimizing actions in isolation without reasoning about future states or long-term goals. This limits…

Artificial Intelligence · Computer Science 2026-03-18 Yongyuan Liang , Shijie Zhou , Yu Gu , Hao Tan , Gang Wu , Franck Dernoncourt , Jihyung Kil , Ryan A. Rossi , Ruiyi Zhang

Even though iterative solvers like the Conjugate Gradients method (CG) have been studied for over fifty years, fault tolerance for such solvers has seen much attention in recent years. For iterative solvers, two major reliable strategies of…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-05-17 Kiril Dichev , Dimitrios S. Nikolopoulos

As distributed machine learning (ML) workloads scale to thousands of GPUs connected by high-speed interconnects, tail latency in collective communication has become a major bottleneck. Existing RDMA transports, such as RoCE, IRN, SRNIC, and…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-30 Ertza Warraich , Ali Imran , Annus Zulfiqar , Shay Vargaftik , Sonia Fahmy , Muhammad Shahbaz

Multi-agent Deep Reinforcement Learning (MADRL) based traffic signal control becomes a popular research topic in recent years. To alleviate the scalability issue of completely centralized RL techniques and the non-stationarity issue of…

Artificial Intelligence · Computer Science 2023-09-08 Hankang Gu , Shangbo Wang , Xiaoguang Ma , Dongyao Jia , Guoqiang Mao , Eng Gee Lim , Cheuk Pong Ryan Wong

Adversarial robustness is a critical measure of a neural network's ability to withstand adversarial attacks at inference time. While robust training techniques have improved defenses against individual $\ell_p$-norm attacks (e.g., $\ell_2$…

Artificial Intelligence · Computer Science 2025-08-26 Ren Wang , Yuxuan Li , Can Chen , Dakuo Wang , Jinjun Xiong , Pin-Yu Chen , Sijia Liu , Mohammad Shahidehpour , Alfred Hero

Recent developments in large language models have sparked interest in efficient pretraining methods. Stagewise training approaches to improve efficiency, like gradual stacking and layer dropping (Reddi et al, 2023; Zhang & He, 2020), have…

Computation and Language · Computer Science 2024-10-15 Abhishek Panigrahi , Nikunj Saunshi , Kaifeng Lyu , Sobhan Miryoosefi , Sashank Reddi , Satyen Kale , Sanjiv Kumar

Applying Machine Learning (ML) techniques to design and optimize computer architectures is a promising research direction. Optimizing the runtime performance of a Network-on-Chip (NoC) necessitates a continuous learning framework. In this…

Networking and Internet Architecture · Computer Science 2019-08-14 Sheng-Chun Kao , Chao-Han Huck Yang , Pin-Yu Chen , Xiaoli Ma , Tushar Krishna

Parallel Continual Learning (PCL) tasks investigate the training methods for continual learning with multi-source input, where data from different tasks are learned as they arrive. PCL offers high training efficiency and is well-suited for…

Machine Learning · Computer Science 2024-07-12 Li Yuepan , Fan Lyu , Yuyang Li , Wei Feng , Guangcan Liu , Fanhua Shang

Environment sensing and fusion via onboard sensors are envisioned to be widely applied in future autonomous driving networks. This paper considers a vehicular system with multiple self-driving vehicles that is assisted by multi-access edge…

Machine Learning · Computer Science 2025-03-26 Xueyao Zhang , Bo Yang , Xuelin Cao , Zhiwen Yu , George C. Alexandropoulos , Yan Zhang , Merouane Debbah , Chau Yuen

Automating network processes without human intervention is crucial for the complex Sixth Generation (6G) environment. Thus, 6G networks must advance beyond basic automation, relying on Artificial Intelligence (AI) and Machine Learning (ML)…

The ongoing transition to renewable energy is increasing the share of fluctuating power sources like wind and solar, raising power grid volatility and making grid operation increasingly complex and costly. In our prior work, we have…

Artificial Intelligence · Computer Science 2023-02-16 Anton R. Fuxjäger , Kristian Kozak , Matthias Dorfer , Patrick M. Blies , Marcel Wasserer

The increase in design cost and complexity have motivated designers to adopt modular design of System on Chip (SoC) by integrating independently designed small chiplets. However, it introduces new challenges for correctness validation,…

Hardware Architecture · Computer Science 2019-10-14 Pritam Majumder , Sungkeun Kim , Jiayi Huang , Ki Hwan Yum , Eun Jung Kim

Efficient multi-robot task allocation (MRTA) is fundamental to various time-sensitive applications such as disaster response, warehouse operations, and construction. This paper tackles a particular class of these problems that we call…

Multiagent Systems · Computer Science 2023-08-21 Steve Paul , Wenyuan Li , Brian Smyth , Yuzhou Chen , Yulia Gel , Souma Chowdhury

Reinforcement Learning (RL) is a pivotal post-training technique for enhancing the reasoning capabilities of Large Language Models (LLMs). However, synchronous RL post-training often suffers from significant GPU underutilization, referred…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-09-26 Wei Gao , Yuheng Zhao , Dakai An , Tianyuan Wu , Lunxi Cao , Shaopan Xiong , Ju Huang , Weixun Wang , Siran Yang , Wenbo Su , Jiamang Wang , Lin Qu , Bo Zheng , Wei Wang

Agentic Reinforcement Learning (RL) enables Large Language Models (LLMs) to perform autonomous decision-making and long-term planning. Unlike standard LLM post-training, agentic RL workloads are highly heterogeneous, combining…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-30 Wei Gao , Yuheng Zhao , Tianyuan Wu , Shaopan Xiong , Weixun Wang , Dakai An , Lunxi Cao , Dilxat Muhtar , Zichen Liu , Haizhou Zhao , Ju Huang , Siran Yang , Yongbin Li , Wenbo Su , Jiamang Wang , Lin Qu , Bo Zheng , Wei Wang

Trajectory planning for multiple robots in shared environments is a challenging problem especially when there is limited communication available or no central entity. In this article, we present Real-time planning using Linear Spatial…

Robotics · Computer Science 2023-04-04 Baskın Şenbaşlar , Wolfgang Hönig , Nora Ayanian

Mobile edge computing (MEC) deployment in a multi-robot cooperation (MRC) system is an effective way to accomplish the tasks in terms of energy consumption and implementation latency. However, the computation and communication resources…

Networking and Internet Architecture · Computer Science 2021-11-23 Rui Yin , Yineng Shen , Huawei Zhu , Xianfu Chen , Celimuge Wu

To efficiently support safety-related vehicular applications, the ultra-reliable and low-latency communication (URLLC) concept has become an indispensable component of vehicular networks (VNETs). Due to the high mobility of VNETs,…

Information Theory · Computer Science 2020-02-19 Haojun Yang , Kan Zheng , Long Zhao , Lajos Hanzo