AdaPID: Adaptive Momentum Gradient Method based on PID Controller for Non-Convex Stochastic Optimization in Deep Learning

Ailun Jian, Xun Li, Weigang Sun, Gaohang Yu · 2025

Momentum-based methods are commonly used in stochastic optimization to accelerate the training of deep learning models. Nevertheless, these methods may cause overshooting, primarily due to the accumulation of stochastic noise over iterations. The PID optimizer addresses this challenge by utilizing gradient difference information; however, it remains limited by the need for manual tuning of learning rates and momentum decay coefficients, constraining its applicability to deep neural networks (DNNs). This paper presents a comprehensive analytical framework for Proportional-Integral-Derivative (PID) based momentum gradient methods from a dynamical systems perspective, demonstrating linear convergence under regularity conditions for nonconvex objectives. Additionally, we introduce an adaptive double-parameter PID optimizer (AdaPID) that automatically adjusts both the learning rate and momentum decay using prior information, while formally establishing convergence rate guarantees for stochastic optimization. Numerical experiments using benchmark functions and real-world datasets demonstrate the improved robustness and performance of the proposed algorithm.

Read the paper · More papers on PaperTik