Distributed Data Parallel Training Based on Cumulative Gradient
Xuebo Zhang, Chengguang Zhang, Min Jiang · 2022
With the development of science and technology, the scale of deep learning models is getting larger and larger. Target detection models trained with a large amount of labeled data can achieve better performance than machine learning models, but large-scale model training is generally slower and requires high Performance workstation support. Based on the Yolov4 target detection model, this paper studies how to quickly shorten the training time of the model through distributed data parallelism. On this basis, a distributed data parallel training method with cumulative gradients is proposed. This method can solve the problem that the input cannot be increased due to the limitation of video memory. For batch problems, experiments prove that this method has a higher speedup and parallel efficiency.