English

Avoiding Cross-Datacenter Collective Congestion via Disaggregated Buffering

Networking and Internet Architecture 2026-05-14 v2

Abstract

LLM training at the scale of tens of thousands of GPUs now spans multiple datacenters (DC), making cross-DC collectives over long-haul links unavoidable. A critical and overlooked bottleneck arises when these collectives collide with intra-DC traffic at the destination - a common pattern in real workloads. The multi-millisecond congestion control loop is too slow to react, triggering severe packet loss and congestion collapse. We present Spillway, a transparent in-network mechanism that buffers dropped packets in switch-disaggregated buffers in a destination data center and drains them once congestion subsides. Through large-scale end-to-end simulations and a hardware prototype, we show that Spillway eliminates performance degradation from collective collisions, reducing iteration time by up to 14 %, without changes to end hosts or training frameworks.

Keywords

Cite

@article{arxiv.2605.11852,
  title  = {Avoiding Cross-Datacenter Collective Congestion via Disaggregated Buffering},
  author = {Mariano Scazzariello and Noga H. Rotman and Dima Gavrilenko and Sajy Khashab and Alexander Shpiner and Matty Kadosh and Marco Chiesa and Dejan Kostic and Mark Silberstein},
  journal= {arXiv preprint arXiv:2605.11852},
  year   = {2026}
}
R2 v1 2026-07-22T07:07:14.590Z