A Method for Parameter Updating in Distributed Deep Learning

Ming He, Xi Guo, Dexuan Wang, Shouhao Qin · 2024

The improvement of distributed training performance of deep learning models involves multiple aspects. This article proposes a parameter update method for distributed training of deep learning models from the perspective of training methods. The method mainly includes two stages of model training. In the first stage, the synchronization period and differential update parameters are mainly obtained; in the second stage, each worker node conducts distributed training based on the synchronization period and differential update parameters obtained in the first stage. This method can reduce the synchronization waiting time between worker nodes, reduce the communication pressure between parameter servers and worker nodes, and improve the performance of distributed training of deep learning models while improving the utilization of computing power resources.

Read the paper · More papers on PaperTik