Layer-wise Top-$k$ Gradient Sparsification for Distributed Deep Learning

Guangyao Li · 2023

Distributed training is widely used in training large-scale deep learning models, and data parallelism is one of the dominant approaches. Data-parallel training has additional communication overhead, which might be the bottleneck of training system. Top-$\boldsymbol{k}$sparsification is a successful technique to reduce the communication volume to break the bottleneck. However, top-$\boldsymbol{k}$sparsification cannot be executed until backpropagation is completed, which disables the overlap of backpropagation computations and gradient communications, leading to limiting the system scaling efficiency. In this paper, we propose a new distributed optimization approach named LKGS-SGD, which combines synchronous SGD (S-SGD) with a novel layer-wise top-$\boldsymbol{k}$sparsification algorithm (LKGS). The LKGS-SGD enables the overlap of computations and communications, and adapts to gradient exchange at layer-wise. Evaluations are conducted by real-world applications. Experimental results show that LKGS-SGD achieves similar convergence to dense S-SGD, while outperforming the original S-SGD and S-SGD with top-$\boldsymbol{k}$sparsification.

Read the paper · More papers on PaperTik