Regularized Over-Parametrized Neural Networks Learned by Gradient Descent Can Generalize Well
Michael Köhler, Adam Krzyżak · IEEE Transactions on Information Theory · 2025
Estimation of univariate regression function by a neural network with one hidden layer is considered, where the weight vector is determined by applying gradient descent to a regularized empirical$L_{2}$risk. Here the number of hidden neurons is allowed to be much larger than the sample size. It is shown that the estimate nevertheless generalizes well in case that the Fourier transform of the regression function decays suitably fast, and that in this case over-parametrization leads to a particular good rate of convergence.