A Memory Enhancement Adjustment Method Based on Stochastic Gradients
Gan Li · 2022 41st Chinese Control Conference (CCC) · 2022
Stochastic gradient descent methods and its variants have been widely used to learn the parameters of a neural network by solving an associated non-convex minimization problem. We propose a new momentum method and adaptive method (MEA/AdaMEA) based on memory enhancement adjustment, which is different from the traditional momentum methods and adaptive methods. At the theoretical level, we prove the convergence of the MEA method for solving a non-convex minimization problem, and then analyze the generalization error of the MEA method from the perspective of consistent stability, showing that the MEA method can improve the stability of the learned model and enhance the generalization performance. At the experimental level, we compare the empirical results of the MEA/AdaMEA method and common optimizers for deep learning on the datasets CIFAR10 and CIFAR100 under the network architectures of the convolutional neural networks ResNet18 and ResNet34. Experiments demonstrate the effectiveness of our proposed MEA/AdaMEA.