Domain-Specific Load Balancing for Accelerating Gradient Synchronization Communication in Large Model Training
Changlong Dong, Qiang Wu, Ran Wang · 2024
As the scale of GPU clusters expands in large model training scenarios, the mismatch between GPU computing speed and gradient data transmission speed within data centers becomes increasingly evident, resulting in gradient synchroni-zation communication becoming a major bottleneck constraining training efficiency. Despite efforts to mitigate this issue through gradient compression techniques for reducing communication volume and scheduling strategies for enhancing computation communication parallelism, these methods have overlooked the pivotal role of load balancing. The traffic patterns encountered during large model training significantly diverge from the typical long-tail distribution found in data centers. Specifically, they exhibit a few large flows in the spatial dimension and high burstiness in the temporal dimension, with the optimization objective focused on minimizing the tail flow completion time (Tail FCT). In response, this paper presents a domain-specific load balancing method tailored for large model training scenarios. This method leverages the Holt-Winters model to predict round-trip times (RTT), enabling anticipation and avoidance of potential congestion paths. Furthermore, it allocates bandwidth resources fairly based on the size of gradient synchronization flows, ensuring overall high synchronization performance. Additionally, the method employs proactive strategies to actively occupy idle paths and flexibly redirect congested flow, thereby adapting to dynamic network environ-ments. Simulation experiments validate that this method significantly outperforms current mainstream load balancing methods under various GPU cluster scales and congestion path percentages.