Data-Induced Multiscale Losses and Efficient Multirate Gradient Descent Schemes
Juncai He, Liangchen Liu, Yen‐Hsi Richard Tsai · Journal of Intelligent Algorithms and Scientific Computing · 2026
This paper investigates the impact of multiscale data, characterized by significant variations in scale across different principal directions, on the optimization dynamics of machine learning algorithms, particularly in deep neural networks. We rigorously establish that such multiscale data induces a corresponding multiscale geometry in the loss landscape. By deriving asymptotic multiscale expansions, we identify explicit component-wise multiscale structure in both the gradients and Hessians. To exploit this structural connection algorithmically, we draw inspiration from multiscale methods in scientific computing and propose Multirate Gradient Descent (MrGD). Rather than relying on empirical learning-rate schedules or being limited by the most restrictive stability constraints, MrGD provides a systematic, data-informed strategy that combines large and small step sizes. Our theoretical and numerical analyses show that this approach goes beyond heuristic tuning, provides rigorous convergence guarantees, and significantly improves the convergence rate.