Accelerated learning in back-propagation nets
J. Schmidhuber · 1989
Two of the most serious problems with back-propagation (bp) (Werbos, 1974)(Parker, 1985)(Rumelhart et al., 1986)(Almeida, 1987) are insufficient speed and the danger of getting stuck in local minima. We offer an approach to cope with both of these problems: Instead of using bp to find zero-points of the gradient of the error-surface we are looking for zero-points of the error-surface itself. This can be done with less computational effort than there is in second order methods. Experimental results indicate that in cases where only a small fraction of units is active simultaneously (sparse coding), this method can be applied successfully. Furthermore it can be significantly faster than conventional bp. Keywords: Back-propagation, sparse coding, speed, learning rate, local minima. 1 The Method Numerous gradient descent methods for adjusting weights in neural nets are described in the literature (see e.g. articles by Parker, Dahl, and Watrous in IEEE 1st Int. Conf. on Neural Networks, V...