The effect of training dynamics on neural network performance

Charles L. Wilson, James L. Blue, Omid M. Omidvar · 1995

In this paper, analysis of a simple model of recurrent network dynamics is used to gain qualitative insights into the training dynamics of multilayer perceptrons (MLPs).These insights allow the training methods used for MLPs to be modified to significantly improve network performance.In previous work [1], the Probabihstic Neural Network (PNN) [2], wasshown to provide better zero-reject error performance on character and fingerprint classifica- tion problems than Radial Basis Function and MLP-based neural network methods.We will show that performance equal to or better than PNN can be achieved with a single three-layer MLP by making fundamental changes in the network optimization strategy.These changes are: 1) Neuron activation functions are used which reduce the probabihty of singular Jacobians; 2) Successive regularization is used to constrain the volume of the minimized weight space; 3) Boltzmann pruning [3] is used to constrain the dimension of the weight space; and 4) Prior class probabilities are used to normalize all error calculations so that statistically significant samples of rare but important classes can be included without distorting the error surface.All four of these changes are made in the inner loop of a conjugate gradient optimization iteration [4] and are intended to simphfy the training dynamics of the optimization.On handprinted digits and fingerprint classification problems these modifications improve error- reject performance by factors between 2 and 4 and reduce network size by 40% to 60%.

Read the paper · More papers on PaperTik