Minimization through convexitization in training neural networks
J.T. Lo · 2003
Provides a mathematical explanation of the ability of the adaptive risk-averting training method to avoid poor local minima. The method actually transforms the standard least-squares error criterion into a "quasi-convex" criterion to make it unnecessary to search throughout the entire weight space to avoid poor local minima. Two theorems are proven in the paper, one examining the convexity region of the risk-averting error criterion to which the standard criterion is transformed to and the other giving a minimax interpretation of the risk-averting error criterion.