A Statistical Test of The Effect of Learning Rate and Momentum Coefficient of Sgd and Its Interaction on Neural Network Performance
Bing-Chuan Chen, Aixiang Chen, Xiaolong Chai, Bian Rui · 2019
Stochastic Gradient Descent (SGD) is a well-received algorithm for large-scale optimization in neural networks for its low iteration cost. However, due to Gradient variance, it often has difficulty in finding optimal learning rate and thus suffers from slow convergence. Using momentum is proven to be a simple effective way of overcoming the slow convergence problem of SDG as long as momentum is properly set. According to the performance metrics, this paper proposes a novel statistical model for analyzing the performance of neural networks. The model takes into account learning rate and momentum, and the method can be used to evaluate and verify their interaction effects on neural network performance. Our study shows that the interaction effects are significant. When momentum has a value smaller than 0.5, the impact on the training time is not statistically noticeable.