Communication-Efficient Distributed Machine Learning: Techniques and Innovations

Juntao Wang · Applied and Computational Engineering · 2025

As Artificial Intelligence (AI) technologies continue to advance, the size and complexity of machine learning models are rapidly increasing. Distributed Machine Learning (DML) has been proposed to improve the limitations of centralized training in terms of computational power and memory. However, communication overhead remains a significant obstacle in DML, restricting training efficiency. This paper proposes a hybrid approach combining Adaptive Gradient Compression (AGC) and Locally Updated Stochastic Gradient Descent (LU-SGD) to maintain model performance while reducing communication overhead. Specifically, the communication load is first reduced by compressing the gradient during each transmission round via AGC. Second, this study uses LU-SGD to minimize the number of communication phases by executing several local updates before synchronizing the gradients. Extensive experiments are conducted on Modified National Institute of Standards and Technology (MNIST), Canadian Institute for Advanced Research (CIFAR)-10, and ImageNet datasets with LeNet, Residual Neural Network (ResNet), and Visual Geometry Group (VGG). Experimental results show the hybrid approach reduces communication data while maintaining efficiency and model accuracy. This approach effectively optimizes communication and demonstrates its potential to improve distributed machine learning frameworks' expansion capability and efficiency.

Read the paper · More papers on PaperTik