An Information-Theoretic View for Deep Learning
Jingwei Zhang, Tongliang Liu, Dacheng Tao · arXiv (Cornell University) · 2018
Deep learning has transformed computer vision, natural language processing, and speech recognition\cite{badrinarayanan2017segnet, dong2016image, ren2017faster, ji20133d}. However, two critical questions remain obscure: (1) why do deep neural networks generalize better than shallow networks; and (2) does it always hold that a deeper network leads to better performance? Specifically, letting $L$ be the number of convolutional and pooling layers in a deep neural network, and $n$ be the size of the training sample, we derive an upper bound on the expected generalization error for this network, i.e., \begin{eqnarray*} \mathbb{E}[R(W)-R_S(W)] \leq \exp{\left(-\frac{L}{2}\log{\frac{1}η}\right)}\sqrt{\frac{2σ^2}{n}I(S,W) } \end{eqnarray*} where $σ>0$ is a constant depending on the loss function, $0