An Adaptive Gradient Method with Differentiation Element in Deep Neural Networks
Runqi Wang, Wei Wang, Teli Ma, Baochang Zhang · 2020
Current adaptive gradient algorithm (such as Adam) used in deep neural network has the advantages of fast training speed, simple tuning task and high computational efficiency. However, these methods are usually based on the gradient update using the root mean square of the past gradient, which often causes the learning rate shock. Thus the model overshoot may be large and even cannot converge. The PID optimization algorithm for deep neural network provides a new way to solve this problem. It introduces the idea of automatic control to solve the problem of overshooting in the stochastic gradient algorithm. The Adam algorithm is similar to an adaptive PI controller. Inspired by this, the differentiation element is introduced into Adam algorithm to accelerate model convergence. The algorithm was tested on MNIST, Cifar-10, Cifar-100 and Tiny-ImageNet data sets in the section of experiment. It is shown that the training speed by 10% on the premise of guaranteeing the accuracy of the model.