SCALE-CCL: A Scalable Collective Communication Library for Wide-Area Distributed Training

Jiaheng Xiong, Qiaolun Zhang, Paolo Medagliani, Michele Ferrero, Xiaomin Liu, Meng Lian, Nicola Di Cicco, Baosen Zhao, Mëmëdhe Ibrahimi, Francesco Musumeci, Massimo Tornatore · 2025

Collective Communication Libraries (CCLs) are widely used to coordinate and optimize data exchange among multiple GPUs. As training clusters increasingly span multiple datacenters for scalability and resource pooling, applying CCLs over wide-area networks (WANs) becomes essential. However, existing CCLs are primarily designed for stable, low-latency intra-datacenter environments. In contrast, WANs exhibit dynamic and unpredictable conditions that limit the effectiveness of traditional CCLs. While recent solutions adopt global optimization techniques, such as Integer Linear Programming (ILP), to improve communication efficiency, these methods assume static topologies and full network visibility, which are often unrealistic in WAN settings with link variability, limited observability, and dynamic traffic patterns. To address these challenges, we propose SCALE-CCL, a scalable heuristic algorithm for optimizing AllGather operations in WAN environments. SCALE-CCL derives global schedules offline from lightweight aggregated metrics (e.g., sampled bandwidth, queue status, and reception history). Thanks to its sub-second synthesis time, schedules can be recomputed whenever needed, either before each training round or when significant WAN changes occur, without noticeable overhead. Compared to the state-of-the-art ILP-based solution TE-CCL and baselines including a shortest-path heuristic (SPH) and NCCL, SCALE-CCL reduces scheduling time by up to four orders of magnitude, while achieving AllGather completion times within a 10% gap of TE-CCL in more than 90% of cases (never exceeding 15%). It consistently outperforms SPH and generally surpasses NCCL in WAN environments.

Read the paper · More papers on PaperTik