Learning and generalization in neural networks
Leon N. Cooper, Charles M. Bachmann · 1990
We review the theoretical basis of backward propagation as a gradient descent algorithm in synaptic weight-space. We then develop an extended version of the model which incorporates gain modification demonstrate that it is equivalent to rescaling the synaptic vectors and replacing the modification step-constant with a time-dependent and direction-dependent step-size. Comparing the extended model with the original model in simulations of a simple two-dimensional paradigm, we find that a combination of gain modification and backward propagation with momentum achieves the most rapid convergence to a close fit to the input-output mapping for the training data. However, the onset of generalization occurs much earlier in training, and on these shorter time scales, ordinary backward propagation with momentum is sufficient to achieve good generalization. Empirical results which we have obtained with stop-consonant tokens from two speech databases demonstrate the ability of backward propagation networks to perform feature extraction. Nevertheless, the generalization accuracy of backprop deteriorates when the dimensionality of the synaptic weight-space becomes too large; this is illustrated in a twelve-class vowel identification paradigm originally studied by Cole, Muthusamy, and Atlas (1990). Working with their data, we show, however, that smaller backprop networks can be trained to isolate pairs of vowels from the other vowel classes to a high degree of accuracy, while other small backprop networks may be used to separate the vowels within a pair. Our results suggest that a hybrid architecture, combining these reduced networks to address the full twelve-class problem, may achieve a higher level of accuracy. To the same end, we also discuss the possibility of combining backward propagation with the RCE neural network (Reilly, Cooper, and Elbaum, 1982). In an appendix, we discuss the limitations of the Hopfield model and contrast them with the improved storage capacity and error-correction properties of a high-density storage model (Bachmann, Cooper, Dembo, and Zeitouni, 1987).