Momentum and optimal stochastic search
Genevieve Orr · 2018
This paper uses the dynamics of weight space probabilities [3, 4] to address stochastic gradient algorithms with learning rate annealing and momentum. This theoretical framework provides a simple, unified treatment of asymptotic convergence rates and asymptotic normality. The results for algorithms without momentum have been previously discussed in the literature. Here we gather those results under a common theoretical structure and extend them to stochastic gradient descent with momentum. DENSITY EVOLUTION AND ASYMPTOTICS