Convergence guarantees for gradient descent in deep neural networks with non-convex loss functions
Agnideep Aich, Agnideep Aich, Ashit Baran Aich, Ashit Baran Aich, Bruce A. Wade · International Journal of Computer Mathematics · 2025
Despite the empirical success of deep neural networks (DNNs), theoretical understanding of their optimization dynamics remains limited due to non-convex loss landscapes. We address this gap by introducing locally quasi-convex regions (LQCRs), regions where gradient descent (GD) exhibits reliable convergence properties. Our key contributions are threefold: (1) A rigorous characterization of LQCRs showing they naturally emerge under standard initialization schemes, (2) Non-asymptotic convergence guarantees for GD with explicit dependence on network depth and width, and (3) A novel depth-aware initialization strategy achieving Θ(L1/2) improvement in convergence rate over conventional methods, resulting in approximately 23% faster convergence in practice. Through both theoretical analysis and experiments on synthetic data and CIFAR-10, we demonstrate that simple gradient methods can provably avoid bad local minima when initialized properly. Our work bridges the gap between neural network practice and convex optimization theory, providing new insights into why gradient-based methods succeed in deep learning.