Deep Residual Learning via Large Sample Mean-Field Stochastic Optimization

Lijun Bo, Agostino Capponi, Huafu Liao · arXiv (Cornell University) · 2019

We study a class of stochastic optimization problems of the mean-field type arising in the optimal training of a deep residual neural network. We estimate the training weights of the network as the optimal relaxed control of a sampling problem, where a population risk criterion is minimized. We establish the existence of optimal relaxed controls when the training set has finite size. The core of our paper is to prove, via $\Gamma$-convergence, that the minimizer of the sampled relaxed problem converges to that of the limiting optimization problem, as the number of training samples grows large. We connect the limit of the sampled objective functional to the unique solution, in the trajectory sense, of a nonlinear Fokker-Planck-Kolmogorov (FPK) equation in a random environment.

Read the paper · More papers on PaperTik