Unified Designs of Multi-Rail-Aware MPI Allreduce and Alltoall Operations Across Diverse GPU and Interconnect Systems

Chen-Chun Chen, Jinghan Yao, Lang Xu, Hari Subramoni, Dhabaleswar K. DK Panda · 2025

The growing demand for computing power in highperformance computing is driving the adoption of diverse accelerators and interconnect networks in modern exascale clusters. Within a node, device interconnects like NVLink, Infinity Fabric, and$\mathbf{X}^{e}$Link, along with IPC techniques, provide high throughput in dense GPU environments. Recently, multi-rail interconnects, such as InfiniBand, Slingshot, and Omni-Path, have enabled highbandwidth communication between nodes. Additionally, highperformance computing applications impose significant demands on collective operations such as Allreduce and Alltoall. Therefore, designing efficient and scalable MPI runtimes for diverse system architectures at large scales is essential. In this paper, we propose unified designs to optimize MPI Allreduce and Alltoall operations using multi-rail-aware, two-level algorithms. The designs support a variety of GPU and interconnect combinations, including NVIDIA, AMD, and Intel GPUs, across InfiniBand, Slingshot, and Omni-Path networks, in the modern dense GPU systems. We optimized the Allreduce operation using a persistent device buffer for device-side reduction and employed an early-triggered, pipelined approach to overlap computation with communication. Additionally, we leveraged the device buffer with IPC techniques as a shared buffer to enhance the two-level Alltoall algorithm, designing both PUSH and PULL variants. We evaluate the advantages of our designs through benchmark and applicationlevel tests on the IsambardAI, Frontier, Cardinal, and Stampede3 systems. In benchmark evaluations, the proposed Allreduce design shows a$2.8 \mathrm{x}, 2.9 \mathrm{x}$, and 2.9 x performance improvement at 1 GB with 32 NVIDIA, 64 AMD, and 64 Intel GPUs, respectively. Additionally, the proposed Alltoall design demonstrates a 1.3 x and 1.05 x improvement at 4 MB with 32 NVIDIA and 64 AMD GPUs, respectively. In application-level evaluations, the proposed Allreduce design demonstrates a$2 x$performance improvement in Amber, while the Alltoall design shows a 1.4x performance gain in heFFTe, both tested on 32 H100 and GH200 GPUs with Infiniband and Slingshot-11 interconnects, respectively.

Read the paper · More papers on PaperTik