Parameter Optimization Algorithms
Chris Bishop · 1995
Abstract In previous chapters, the problem of learning in neural networks has been formulated in terms of the minimization of an error function E. This error is a function of the adaptive parameters (weights and biases) in the network, which we can conveniently group together into a single W-dimensional weight vector w with components. In Chapter 4 it was shown that, for a multi-layer perceptron, the derivatives of an error function with respect to the network parameters can be obtained in a computationally efficient way using back-propagation. We shall see that the use of such gradient information is of central importance in finding algorithms for network training which are sufficiently fast to be of practical use for large-scale applications.