ADTopk: All-Dimension Top-k Compression for High-Performance Data-Parallel DNN Training

Zhangqiang Ming, Yuchong Hu, Wenxiang Zhou, Xin Zheng, Chenxuan Yao, Dan Feng · 2024

Data-parallel deep neural networks (DNN) training systems deployed across nodes have been widely used in various domains, while the system performance is often bottlenecked by the communication overhead among workers for synchronizing gradients. Top-k sparsification compression is the de facto approach to alleviate the communication bottleneck, which truncates the gradient to its largest k elements before sending it to other nodes.

Read the paper · More papers on PaperTik