Towards Understanding Variation-Constrained Deep Neural Networks
Gen Li, Jie Ding · IEEE Transactions on Signal Processing · 2023
Multi-layer feedforward networks have been used to approximate a wide range of nonlinear functions. A fundamental problem is understanding the generalizability of a neural network model through its statistical risk, or the expected test error. In particular, it is important to understand the phenomenon that overparameterized deep neural networks may not suffer from overfitting when the number of neurons and learning parameters rapidly grow with$n$or even surpass$n$. In this paper, we show that a class of variation-constrained regression neural networks, with arbitrary width, can achieve a near-parametric rate$n^{-1/2+\delta }$for an arbitrarily small positive constant$\delta$. It is equivalent to$n^{-1 +2\delta }$under the mean squared error. This rate is also observed from numerical experiments. The result provides an insight into the benign overparameterization phenomenon. It indicates that the number of trainable parameters may not be a suitable complexity measure as often perceived for classical regression models. We also discuss the convergence rate regarding other network parameters, including the input dimension, network layer, and coefficient norm.