TCCL: Co-optimizing Collective Communication and Traffic Routing for GPU-centric Clusters

Baojia Li, Xiaoliang Wang, Jingzhu Wang, Yifan Liu, Yuanyuan Gong, Hao Lu, Weizhen Dang, Weifeng Zhang, Xiaojie Huang, Mingzhuo Chen, Jie Chen, Chunzhi He, Yadong Liu, Xiaoyuan Hu, Chen Liu, Xuefeng Ji, Yinben Xia, Xiang Li, Zekun He, Yachen Wang · 2024

GPU-centric clusters are increasingly deployed to support many AI services including the task of large language model (LLM) training. Notably, the corresponding networks have demonstrated multiple new characteristics, such as the boundary of the network operation has been extended from switch network to GPU interconnection network; the communication pattern is regular and predictable; the training is easily affected by the network jitters. These introduce new challenges and opportunities for network operators to build high-performance resilient networking systems. In this paper, we present TCCL, an operational practice to manage GPU-centric networks with over 10K heterogeneous GPU cards. We argue that GPU-centric networks require joint optimization of topology-aware collective communication at the host and centralized routing management in the multi-path network. By leveraging the characteristics of the GPU-centric network, TCCL fully utilizes the high bandwidth of both GPU and switch networks in parallel, reduces the delay of collective communication with short paths across nodes, and avoids network congestion caused by route conflict through traffic planning.

Read the paper · More papers on PaperTik