CoLa: Towards Communication-efficient Distributed Sparse Matrix-Matrix Multiplication on GPUs

Lixing Zhang, Yingxia Shao, Shigang Li · 2025

Sparse Matrix-Matrix Multiplication (SpMM) is a critical operator in many applications, such as graph neural networks (GNNs).However, when SpMM is scaled to multiple GPUs, existing works face significant challenges: (1) massive redundant communication and (2) unawareness of heterogeneous links.To address the issue of massive redundant communication, we introduce a communication redundancy-free distributed SpMM algorithm that efficiently reutilizes fetched remote data to reduce communication volume.To tackle the unawareness of heterogeneous links, we propose two link-aware optimization techniques: communication fusion, which leverages local GPUs as embedding caches to reduce communication over slow links; and requestcoalesced communication, which coalesces necessary requested remote data into bulk transfers to maximize bandwidth utilization and minimize communication volume over proxy-based links.Based on these techniques, we develop CoLa, a highly communication-efficient distributed SpMM framework.Extensive evaluations on real-world datasets under different multi-GPU settings demonstrate that CoLa achieves geomean speedups of 8.56×, 9.12×, and 57.97× over CAGNET, MGGCN, and MGG, respectively.

Read the paper · More papers on PaperTik