Sparse Collectives: Exploiting Data Sparsity to Improve Communication Efficiency
Dhananjaya Wijerathne, Haris Javaid, Guanwen Zhong, Dan Wu, Xing Yuan Kom, Mario Baldi · 2025
The scalability of distributed Machine Learning (ML) systems heavily relies on efficient data exchange between multiple devices, which is typically achieved through collective operations such as all-gather, reduce-scatter and allreduce. Many ML models, including transformer-based models, exhibit significant sparsity in activations and gradients that are exchanged through these collectives. In this work, we introduce lightweight sparse collectives designed to exploit this data sparsity, aiming to minimize the communication volume while keeping the overhead of sparsity-aware compression and decompression low. These sparse collectives deliver substantial improvements in collective performance, significantly reducing their completion time. Experiments on AMD Instinct™ MI210 and MI300X GPU nodes demonstrate up to 2.96× allreduce, 2.6× allgather, and 2.85× reduce-scatter speedup. These results highlight the potential of sparse collectives to accelerate large-scale distributed training and inference systems.