English

HiCCL: A Hierarchical Collective Communication Library

Distributed, Parallel, and Cluster Computing 2024-08-13 v1

Abstract

HiCCL (Hierarchical Collective Communication Library) addresses the growing complexity and diversity in high-performance network architectures. As GPU systems have envolved into networks of GPUs with different multilevel communication hierarchies, optimizing each collective function for a specific system has become a challenging task. Consequently, many collective libraries struggle to adapt to different hardware and software, especially across systems from different vendors. HiCCL's library design decouples the collective communication logic from network-specific optimizations through a compositional API. The communication logic is composed using multicast, reduction, and fence primitives, which are then factorized for a specified network hieararchy using only point-to-point operations within a level. Finally, striping and pipelining optimizations applied as specified for streamlining the execution. Performance evaluation of HiCCL across four different machines\unicodex2014\unicode{x2014}two with Nvidia GPUs, one with AMD GPUs, and one with Intel GPUs\unicodex2014\unicode{x2014}demonstrates an average 17×\times higher throughput than the collectives of highly specialized GPU-aware MPI implementations, and competitive throughput with those of vendor-specific libraries (NCCL, RCCL, and OneCCL), while providing portability across all four machines.

Keywords

Cite

@article{arxiv.2408.05962,
  title  = {HiCCL: A Hierarchical Collective Communication Library},
  author = {Mert Hidayetoglu and Simon Garcia de Gonzalo and Elliott Slaughter and Pinku Surana and Wen-mei Hwu and William Gropp and Alex Aiken},
  journal= {arXiv preprint arXiv:2408.05962},
  year   = {2024}
}
R2 v1 2026-06-28T18:10:08.496Z