Quasi-Adam: Accelerating Adam Using Quasi-Newton Approximations

Aditya Ranganath, Irabiel Romero Ruiz, Mukesh Kumar Singhal, Roummel F. Marcia · 2024

Adam is arguably one of the most commonly used approach in deep learning and machine learning. With good regret bounds and empirical convergence proofs, the approach has produced many state-of-the-art models over a variety of problems. However, the method only uses gradient information at each iterate in addition to some moving averaged gradients and its corresponding moments from the past. In this paper, we propose a method that builds upon Adam and incorporates quasi-Newton matrices for approximating second derivatives. These Hessian approximations satisfy the so-called secant equation, which is the first-order Taylor series expansion of the gradient along the direction of the change in iterates. Judicious choices of quasi-Newton matrices can lead to guaranteed descent in the objective function and improved convergence. In this work, we integrate search directions obtained from using these quasi-Newton Hessian approximations with the Adam optimization algorithm. We provide convergence guarantees and demonstrate improved performance through an extensive experimentation on a variety of applications.

Read the paper · More papers on PaperTik