English

FAST: An Efficient Scheduler for All-to-All GPU Communication

Distributed, Parallel, and Cluster Computing 2026-03-09 v3 Networking and Internet Architecture

Abstract

All-to-All(v) communication is a critical primitive in modern machine learning workloads, particularly mixture-of-experts (MoE) models. Unfortunately, efficient scheduling is challenging due to workload skew, heterogeneous two-tier fabrics, and incast congestion, compounded by the dynamic nature of MoE workloads, where traffic shifts every few hundred milliseconds. Existing schedulers are hardly scalable, incurring seconds to hours of synthesis time, making them impractical. We present FAST, an efficient All-to-All(v) scheduler. FAST addresses skew through intra-server rebalancing and enforces balanced, one-to-one scale-out transfers that avoid incast. Evaluated extensively on both NVIDIA H200 and AMD MI300X clusters, FAST consistently outperforms state-of-the-art solutions on skewed workloads while reducing synthesis time by orders of magnitude.

Keywords

Cite

@article{arxiv.2505.09764,
  title  = {FAST: An Efficient Scheduler for All-to-All GPU Communication},
  author = {Yiran Lei and Dongjoo Lee and Liangyu Zhao and Daniar Kurniawan and Chanmyeong Kim and Heetaek Jeong and Changsu Kim and Hyeonseong Choi and Liangcheng Yu and Arvind Krishnamurthy and Justine Sherry and Eriko Nurvitadhi},
  journal= {arXiv preprint arXiv:2505.09764},
  year   = {2026}
}

Comments

Accepted to 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 2026)

R2 v1 2026-06-28T23:33:39.922Z