Accelerating Distributed Training through In-network Aggregation with Idle Resources in Data Centers
Ge Wang, Huaqing Tu, Gongming Zhao, Hongli Xu · 2023
As the parameter scale of large-scale models continues to increase, distributed model training imposes significant communication overhead in data centers, resulting in reduced training efficiency. To address this challenge, a promising solution, called in-network aggregation, is proposed to mitigate the bandwidth bottleneck in data centers by aggregating gradients within network. However, existing works mainly rely on programmable switches to implement in-network aggregation, while programmable switches have limited memory resources, which makes it hard to storage parameter with large size. Moreover, programmable switches are not yet widely deployed in current data centers, resulting in poor availability. To overcome this limitation, we leverage servers with idle resources in data centers for in-network aggregation, since servers have more powerful storage capabilities compared with programmable switches. Specifically, we formally formulate the problem of server-based in-network aggregation. An efficient approximate algorithm with bounded approximation factor is proposed to select servers with idle resources and paths for model aggregation. Our extensive simulations show that our proposed method can reduce communication time by 38.4%-60.1% compared to state-of-the-art solutions.