Global Minima of Overparameterized Neural Networks

Yaim Cooper · SIAM Journal on Mathematics of Data Science · 2021

We explore some mathematical features of the loss landscape of overparameterized neural networks. A priori, one might imagine that the loss function looks like a typical function from $\mathbb{R}^d$ to $\mathbb{R}$, in particular, that it has discrete global minima. In this paper, we prove that in at least one important way, the loss function of an overparameterized neural network does not look like a typical function. If a neural net has $d$ parameters and is trained on $n$ data points $(x_i, y_i) \in \mathbb{R}^s \times \mathbb{R}^r$, with $d>r n$, we show that the locus $M$ of global minima of $L$ is usually not discrete but rather an $(d- rn)$-dimensional submanifold of $\mathbb{R}^d$. In practice, neural nets commonly have orders of magnitude more parameters than data points, so this observation implies that $M$ is typically a very-high-dimensional submanifold of $\mathbb{R}^d$.

Read the paper · More papers on PaperTik