Parallel Gradient Computation and Synchronization: Enhancing the Efficiency of Distributed Training for LLMs
Hao Li, Hao Jiang, Jing Wu, Guiao Yang, Jian Zhang · IEEE Transactions on Network Science and Engineering · 2025
As the size of large language models (LLMs) increases, the limitations of a single data center, such as constrained computational resources and storage capacity, have made distributed training across multiple data centers the preferred solution. However, a primary challenge in this context is reducing the impact of gradient synchronization on the training efficiency across multiple data centers. In this work, we propose a distributed training scheme for LLMs, named parallel gradient computation and synchronization (PGCS). Specifically, while one expert model is being trained to compute gradients, another expert model performs gradient synchronization in parallel. In addition, a gradient synchronization algorithm named BLP is developed to find the optimal gradient synchronization strategy under arbitrary network connectivity and limited bandwidth across multiple data centers. Ultimately, the effectiveness of PGCS and BLP in enhancing the efficiency of distributed training is demonstrated through comprehensive simulations and physical experiments.