A Multidimensional Communication Scheduling Method for Hybrid Parallel DNN Training

Shengwei Li, Kai Lü, Zhiquan Lai, Weijie Liu, Keshi Ge, Dongsheng Li · IEEE Transactions on Parallel and Distributed Systems · 2024

The transformer-based deep neural network (DNN) models have shown considerable success across diverse tasks, prompting widespread adoption of distributed training methods such as data parallelism and pipeline parallelism. With the increasing parameter number, hybrid parallel training becomes imperative to scale training. The primary bottleneck in scaling remains the communication overhead. The communication scheduling technique, emphasizing the overlap of communication with computation, has demonstrated its benefits in scaling. However, most existing works focus on data parallelism, overlooking the nuances of hybrid parallel training. In this paper, we proposeTriRace, an efficient communication scheduling framework for accelerating communications in hybrid parallel training of asynchronous pipeline parallelism and data parallelism. To achieve effective computation-communication overlap,TriRaceintroduces3D communication scheduling, which adeptly leverages data dependencies between communication and computations, efficiently scheduling AllReduce communication, sparse communication, and peer-to-peer communication in hybrid parallel training. To avoid possible communication contentions,TriRacealso incorporates atopology-aware runtimewhich optimizes the execution of communication operations by considering ongoing communication operations and real-time network status. We have implemented a prototype ofTriRacebased on PyTorch and Pipedream-2BW, and conducted comprehensive evaluations with three representative baselines. Experimental results show thatTriRaceachieves up to 1.07–1.45× speedup compared to the state-of-the-art pipeline parallelism training baseline Pipedream-2BW, and 1.24–1.81× speedup compared to the Megatron.

Read the paper · More papers on PaperTik