Generalization Performance of Empirical Risk Minimization on Over-Parameterized Deep ReLU Nets

Shao-Bo Lin, Yao Wang, Ding‐Xuan Zhou · IEEE Transactions on Information Theory · 2025

In this paper, we study the generalization performance of global minima of empirical risk minimization (ERM) on over-parameterized deep ReLU nets. Using a novel deepening scheme for deep ReLU nets, we rigorously prove that there exist perfect global minima achieving optimal generalization error rates for numerous types of data under mild conditions. Since over-parameterization of deep ReLU nets is crucial to guarantee that the global minima of ERM can be realized by the widely used stochastic gradient descent (SGD) algorithm, our results present a potential way to fill the gap between optimization and generalization of deep learning.

Read the paper · More papers on PaperTik