Natasha 2: Faster Non-Convex Optimization Than SGD

Zeyuan Allen-Zhu · Neural Information Processing Systems · 2018

(this is a theory paper) We design a stochastic algorithm to find e -approximate local minima of any smooth nonconvex function in rate O(e−3.25) , with only oracle access to stochastic gradients. The best result was essentially O(e−4) by stochastic gradient descent (SGD).

Read the paper · More papers on PaperTik