Learning Shallow Neural Networks via Provable Gradient Descent with Random Initialization

Shuhao Xia, Yuanming Shi · 2019

This paper presents the provable gradient descent algorithm with random initialization for learning a two-layer neural network with quadratic activation functions. Specifically, we focus on the under-parameterized regime where the number of hidden units is smaller than the dimension of the inputs. We reveal that the randomly initialized gradient descent for the nonconvex neural network training problem is able to enter a local region that enjoys strong convexity and strong smoothness within a few iterations, and then provably converges to a globally optimal model at a linear rate.

Read the paper · More papers on PaperTik