A Survey of Neural Network Optimization Algorithms

Chengwei Ji · 2024

This paper aims to explore seven commonly used optimization algorithms in deep learning: SGD, Momentum-SGD, NAG, AdaGrad, RMSprop, AdaDelta, and Adam. Based on an overview of their theories and development histories, this paper constructs convolutional neural networks and BERT to conduct numerical experiments on these optimization algorithms. The experimental results show that among the seven optimization algorithms, SGD may get stuck in saddle points, while the other six algorithms developed based on SGD have the ability to escape from saddle points. Compared to other algorithms, Adam converges the fastest and the parameters obtained perform better in practice.

Read the paper · More papers on PaperTik