Sparse Gradient Communication with AlltoAll for Accelerating Distributed Deep Learning
Jing Peng, Zihan Li, Shaohuai Shi, Bo Li · 2024
Synchronous stochastic gradient descent (S-SGD) with data parallelism has become a de-facto approach in training large-scale deep neural networks (DNNs) on multi-GPU systems. However, S-SGD requires iteratively synchronizing gradients from all workers, which incurs excessive communication costs and limits the scaling efficiency of GPU clusters. Gradient sparsification like top-k sparsification techniques has been shown to be potentially effective in reducing the communication volume. Yet, existing approaches with top-k sparsification still suffer from high communication complexity in that they often need a very low density to achieve better performance while easily sacrificing the model accuracy. To this end, we propose a novel sparse communication approach called TopKA2A, which integrates top-k sparsification with AlltoAll communication to exchange sparse tensors among GPUs, resulting in a notable reduction in communication complexity. With rigorous theoretical analysis on the conditions that TopKA2A can be applied, we design a simple yet effective tensor fusion algorithm based on binary search. We perform in-depth analysis and evaluation of communication efficiency by comparing TopKA2A with state-of-the-art solutions over popular models without compromising model accuracy. Experimental results demonstrate that TopKA2A yields substantial communication efficiency gains and runs up to 73% faster than existing algorithms on a 32-GPU cluster.