The Analysis of the Neural Network Optimizers in Condition of the Limited Dataset
A.V. Pchelin A.V. Pchelin, Andrey Sergeevich Martyanov, Dmitry S. Antipin · 2024
Training the neural networks on an ever-smaller amount of data is getting more relevant today. Especially the business is being involved into the tending to spend less computing resources, at the same time reducing the development time of the final product. In this case it is evident that the end user would get the product for less cost in a shorter time. The solution is relatively complex and affects several aspects of neural network architecture development at the same time. One of the key elements of neural network training is the weighting factor optimizer. In this paper the behavior of four popular neural network optimizers - stochastic gradient descent (SGD), root-mean-square propagation (RMSProp), adaptive moment estimation (Adam) and decoupled weight decay regularization (AdamW) were studied for the limited size of training dataset. Each of the optimization methods is described and analyzed in details. As a neural network, a residual neural network (ResNet) architecture has been developed for classifying as many as ten classes. Stochastic gradient descent algorithm turned out to be the best optimizer in terms of accuracy. Root-mean-square propagation algorithm appeared to be the best optimizer related to the loss function. The results of the study are illustrated by graphs of neural network loss functions. The results of this research could be used while training the neural networks with a limited or no access to big data sets.