DynPipe: Toward Dynamic End-to-End Pipeline Parallelism for Interference-Aware DNN Training
Zhengyi Yuan, Xiong Wang, Yuntao Nie, Yufei Tao, Yuqing Li, Zhiyuan Shao, Xiaofei Liao, Bo Li, Hai Jin · IEEE Transactions on Parallel and Distributed Systems · 2025
Pipeline parallelism has emerged as an indispensable technique for training large deep neural networks. While existing asynchronous pipeline systems address the time bubbles inherent in synchronous architectures, they continue to suffer frominefficiencyandsusceptibilitytovolatilehardware environment due to their suboptimal andstaticconfigurations. In this paper, we propose DynPipe, aninterference-awareasynchronous pipeline framework to optimize theend-to-endtraining performance in highlydynamiccomputing environments. By characterizing thenon-overlappedcommunication overheads andconvergencerate conditioned on stage-wise staleness, DynPipe carefully crafts an optimized pipeline partition that harmonizes the hardware speed with statistical convergence. Moreover, DynPipe deploys anon-intrusiverandom forest model that utilizes runtime stage statistics to evaluate the impact of environmental changes, such as task interference and network jitter, on the training efficiency. Following the evaluation guidance, DynPipe adaptivelyadjustspartition plan to restore both intra and inter-stage load balancing, thereby facilitating seamless pipeline reconfiguration in dynamic environments. Extensive experiments show that DynPipe outperforms state-of-the-art systems, accelerating the time-to-accuracy by1.5-3.4×.