English

Scaling All-to-all Operations Across Emerging Many-Core Supercomputers

Distributed, Parallel, and Cluster Computing 2026-01-27 v1

Abstract

Performant all-to-all collective operations in MPI are critical to fast Fourier transforms, transposition, and machine learning applications. There are many existing implementations for all-to-all exchanges on emerging systems, with the achieved performance dependent on many factors, including message size, process count, architecture, and parallel system partition. This paper presents novel all-to-all algorithms for emerging many-core systems. Further, the paper presents a performance analysis against existing algorithms and system MPI, with novel algorithms achieving up to 3x speedup over system MPI at 32 nodes of state-of-the-art Sapphire Rapids systems.

Keywords

Cite

@article{arxiv.2601.17606,
  title  = {Scaling All-to-all Operations Across Emerging Many-Core Supercomputers},
  author = {Shannon Kinkead and Jackson Wesley and Whit Schonbein and David DeBonis and Matthew G. F. Dosanjh and Amanda Bienz},
  journal= {arXiv preprint arXiv:2601.17606},
  year   = {2026}
}
R2 v1 2026-07-01T09:18:47.947Z