A superior error function for training neural networks
Barry L. Kalman, Stan C. Kwasny · 2002
The authors present an error function 'kerr' which does not have the objectionable properties exhibited by the error function 'ferr' usually used to train neural networks. When combined with the conjugate gradient method, 'kerr' shows dramatic speedups over 'ferr' and back-propagation. Ferr is a sum of squares; in kerr each square is divided by a factor.>