Dynamics of On-Line Gradient Descent Learning for Multilayer Neural Networks
David Saad, Sara A. Solla · Aston Publications Explorer (Aston University) · 1995
We consider the problem of on-line gradient descent learning for general two-layer neural networks. An analytic solution is presented and used to investigate the role of the learning rate in controlling the evolution and convergence of the learning process. Learning in layered neural networks refers to the modification of internal parameters fJg which specify the strength of the interneuron couplings, so as to bring the map f J implemented by the network as close as possible to a desired map ~ f . The degree of success is monitored through the generalization error, a measure of the dissimilarity between f J and ~ f . Consider maps from an N-dimensional input space ¸ onto a scalar i, as arise in the formulation of classification and regression tasks. Two-layer networks with an arbitrary number of hidden units have been shown to be universal approximators [1] for such N-to-one dimensional maps. Information about the desired map ~ f is provided through independent examples (¸ ¯ ; ...