Improved generalization and network pruning using adaptive Laplace regularization
Peter M. Williams · 1993
Neural networks designed for regression or classification need to be trained using some form of stabilization or regularization if they are to generalize well beyond the original training set. This means finding a balance between complexity of the network and information content of the data. This paper examines a type of formal regularization in which the penalty term is proportional to the logarithm of the L/sub p/ norm of the weight vector log ( Sigma /sub j/ mod omega /sub j/ mod /sup p/)/sup 1/p/ (p>or=1). The specific choice p=1 simultaneously provides both forms of stabilization with radical pruning leading to greatly improved generalization.