Analysis and Synthesis of Adaptive Gradient Algorithms in Machine Learning: The Case of AdaBound and MAdamSSM

Kushal Chakrabarti, Nikhil Chopra · 2022 IEEE 61st Conference on Decision and Control (CDC) · 2022

Adaptive gradient algorithms have become the prevalent tool in training complex neural networks; recent examples include AdaBound and MAdam. For a better understanding of such existing optimization algorithms and design ideas for new algorithms, well-known tools from classical control appear to be promising. This area of research is built upon modeling optimization algorithms as closed-loop dynamical systems. Consequently, this paper exploits a control-theoretic methodology in analyzing AdaBound, a recent adaptive gradient algorithm, and proposing a novel optimizer for machine learning. Specifically, inspired by the recently developed state-space perspective in the G-AdaGrad and the Adam algorithm, we present a simple convergence analysis of the AdaBound algorithm for non-convex optimization problems. Next, we propose a new variant of the MAdam algorithm upon applying the concept of transfer functions. Our experimental results demonstrate the efficiency of the proposed algorithm in training CNN models for image classification problems. The findings in this paper suggest further exploration of the existing tools from control theory in complex machine learning problems.

Read the paper · More papers on PaperTik