Adathm: Adaptive Gradient Method Based on Estimates of Third-Order Moments
Huikang Sun, Lize Gu, Bin Sun · 2019
Deep learning has been widely used in the field of data aggregation and fusion. As a significant aspect of deep learning, stochastic optimization algorithm affects the operating efficiency and final effect. Adaptive optimization methods such as Adagrad, RMSprop, Adam, which have been proposed to achieve a rapid training process with an element-wise scaling term on learning rates. Nevertheless, unstable and extreme learning rates may fail to converge to an optimal solution (or a critical point in nonconvex settings). In order to reduce the impact of unsuitable learning rate, we proposed a new method that we apply third-order moments to Adam, which is called Adathm. We also introduce the ideal of dynamic bounds on learning rates and endow the proposed method with "long-term memory" of past gradients. Our preliminary experimental results show that our proposed algorithm can fix the convergence issues and compares favorably to other stochastic optimization methods in some real applications.