Analysis of the Impact of the Batch Size and Worker Number in Data-Parallel Training of CNNs
Jiajie Shi, Yueyue Huang, Xinyuan Zhou, Yimeng Xu · 2024
Deep learning has demonstrated robust performance and wide-ranging application prospects across various fields. However, deploying deep learning models often demands substan-tial time and computational resources for training. This study investigates the impact of batch size and number of worker processes on the training time of convolutional neural network (CNN) models. The research employs a dynamic adjustment approach based on gradients and loss to optimize batch sizes, examining the specific effects of different worker process counts on the training efficiency of CNN models such as AlexNet, ResNet, and LeNet. The results reveal significant influences of worker process count selection on training efficiency for the AlexN et model, particularly showing improved performance with 3 and 7 worker processes when batch size is set to 128. For ResN et models, distinct trends in training efficiency are observed across different batch sizes. LeNet models exhibit a tendency towards increasing convergence loss values with higher batch sizes, with the impact of worker process count on training time showing complex and varied effects. This study analyzed the impact of batch size and number of workers on training efficiency when using the MNIST dataset for CNN model training. The experimental results indicate that changes in hyperparameters have a significant impact on the training efficiency of different models. However, as this study only used the MNIST dataset, it is not yet possible to determine whether these findings are applicable to other datasets. Therefore, further research is needed in the future to validate the universality of these phenomena. This study aims to provide universal strategies for parameter selection in various CNN models and offer valuable insights to practitioners, thereby reducing trial-and-error costs and acceler-ating the model development process. These findings contribute theoretical foundations and practical guidance for optimizing and deploying deep learning models. Here is the link to the open-source code repository for researchers to reproduce the results: https://github.com/XIAOYAOGVAGVA/Analysis-of-the-Impact-of-the-Batch-Size-and-Worker-Number-in-Data-Parallel-Training-of-CNNs