Katyusha: the first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2017
Nesterov's momentum trick is famously known for accelerating gradient descent, and has been proven useful in building fast iterative algorithms. However, in the stochastic setting, counterexamples exist and prevent Nesterov's momentum from providing similar acceleration, even if the underlying problem is convex.