CommBench: Micro-Benchmarking Hierarchical Networks with Multi-GPU, Multi-NIC Nodes
Mert Hidayetoğlu, Simon Garcia de Gonzalo, Elliott Slaughter, Yu Li, Christopher J. Zimmer, Tekin Biçer, Bin Ren, William Gropp, Wen‐mei Hwu, Alex Aiken · 2024
Modern high-performance computing systems have multiple GPUs and network interface cards (NICs) per node. The resulting network architectures have multilevel hierarchies of subnetworks with different interconnect and software technologies. These systems offer multiple vendor-provided communication capabilities and library implementations (IPC, MPI, NCCL, RCCL, OneCCL) with APIs providing varying levels of performance across the different levels. Understanding this performance is currently difficult because of the wide range of architectures and programming models (CUDA, HIP, OneAPI).