CE-SGD: Communication-Efficient Distributed Machine Learning
Zeyi Tao, Qi Xia, Qun Li, Songqing Cheng · 2021 IEEE Global Communications Conference (GLOBECOM) · 2021
Training large-scale machine learning models usually demands a distributed approach to process the huge amount of training data efficiently. However, the high network communication cost introduced by parallel stochastic gradient descent (SGD) algorithms is a well-known bottleneck. To this end, we propose CE-SGD, a communication-efficient distributed machine learning algorithm that aggressively reduces the amount of gradient data exchanged among the training workers. CE-SGD belongs to the family of gradient sparsification schemes. CE-SGD adaptively adjusts the gradient sparsity according to the model's feedback. It also selectively transmits the gradients based on their degree of participation in the backpropagation. We mathematically prove the convergence of CE-SGD for both convex and non-convex cases and conduct a series of experiments on our CE-SGD implementation. Our experiments reveal that CE-SGD can achieve fast convergence, desirable gradient compression ratio, and high accuracy with low network bandwidth cost compared to state-of-the-art algorithms.