English

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods

Distributed, Parallel, and Cluster Computing 2026-07-07 v1 Artificial Intelligence Networking and Internet Architecture

Abstract

The deployment of Mixture-of-Experts (MoE) models on production high-bandwidth superpods, such as NVIDIA's NVL72/576 and Huawei's CloudMatrix384, introduces critical challenges beyond raw interconnect bandwidth. While these systems provide unified global address spaces and high-bandwidth fabrics, their full potential for sparse MoE communication is hindered by three fundamental bottlenecks: (1) Strict execution serialization imposed by coarse-grained Bulk Synchronous Parallel (BSP) orchestration of interdependent communication phases; (2) Prohibitive synchronization overhead that fails to scale alongside high interconnect bandwidth; and (3) Severe load imbalance resulting from distance-agnostic scheduling of irregular token traffic. To eliminate these bottlenecks, we introduce UBEP (Unified-Bus Expert Parallelism), a production-ready communication library that rethinks MoE's All-to-All primitives for modern superpod architectures. Through large scale experiments, UBEP reduces All-to-All latency by up to 52.4% and MoE inference Time Per Output Token (TPOT) by up to 11.1%.

Keywords

Cite

@article{arxiv.2607.06202,
  title  = {UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods},
  author = {Yipeng Liu and Chang Liu and Si Shen and Jiaqi Zheng and Mingfan Li and Yuyang Yang and Guanhua Li and Yuquan Zhang and Yimeng Xu and Zhongzhe Hu and Zhiyuan Huang and Qihang Duan and Junsong Wang and Wenkai Ling and Baochuan Yang and Xianzhi Yu and Han Bao and Yijie Chen and Guihai Chen},
  journal= {arXiv preprint arXiv:2607.06202},
  year   = {2026}
}