Analysis of the Rate of Convergence of an Over-Parametrized Deep Neural Network Estimate Learned by Gradient Descent
Michael Köhler, Adam Krzyżak · IEEE Transactions on Information Theory · 2025
Estimation of a regression function from independent and identically distributed random variables is considered. The$L_{2}$error with integration with respect to the design measure is used as an error criterion. Over-parametrized deep neural network estimates are defined which are based on a special network topology, which use a special random initialization and where all the weights are learned by the gradient descent. It is shown that the expected$L_{2}$error of these estimates converges to zero with the rate close to$n^{-1/(1+d)}$in case that the regression function is Hölder smooth with Hölder exponent$p \in [{1/2,1}]$. In case of an interaction model where the regression function is assumed to be a sum of Hölder smooth functions where each of the functions depends only on$d^{*}$of ofdcomponents of the design variable, it is shown that these estimates achieve the corresponding$d^{*}$-dimensional rate of convergence.