Communication-Efficient Distributed LLM Training with Adaptive Traffic Control

Haoying Jin · 2024

The emergence of large language models (LLMs) calls for efficient distributed training approaches for machine learning models. In the traditional PS (Parameter Server)-based distributed LLM training architecture, one or more servers send model parameters to multiple workers, who perform model training using local data and send back gradients to servers to update model weights until the model converges. However, due to the large gradients produced by workers, the server faces high communication pressure caused by the gradients sent from the network (usually from a switch). Aiming to solve this issue, this paper first analyzes the challenges of distributed LLM training and summarizes existing solutions proposed in recent years. Then, we propose a new architecture based on adaptive traffic control, SwitchATC, to reduce the communications between the switch and the PS server. SwitchATC monitors the traffic rate between the switch and the server. If the traffic rate exceeds the server’s network bandwidth, the switch will remove sparse gradients to reduce the communication cost. After introducing the architecture of SwitchATC, the technological challenges of implementing SwitchATC are discussed.

Read the paper · More papers on PaperTik