On-line learning with adaptive back-propagation in two-layer networks
Ansgar H. L. West, David Saad · Physical review. E, Statistical physics, plasmas, fluids, and related interdisciplinary topics · 1997
An adaptive back-propagation algorithm parametrized by an inverse temperature $\ensuremath{\beta}$ is studied and compared with gradient descent (standard back-propagation) for on-line learning in two-layer neural networks with an arbitrary number of hidden units. Within a statistical mechanics framework, we analyze these learning algorithms in both the symmetric and the convergence phase for finite learning rates in the case of uncorrelated teachers of similar but arbitrary length $T$. These analyses show that adaptive back-propagation results generally in faster training by breaking the symmetry between hidden units more efficiently and by providing faster convergence to optimal generalization than gradient descent.