An equivalence between sigmoidal gain scaling and training with noisy (jittered) input data
Russell Reed, Robert J. Marks, Sung‐Kwun Oh · 1992
Training with additive input noise (jitter) in a commonly used heuristic for improving generalization in layered perceptron artificial neural networks. A drawback of training with jitter, in comparison with the unjittered case, is that many more sample presentations are required in order to average over the noise and estimate the expected response. The authors demonstrate that the expected effect of jitter can be computed, in certain cases, by a simple scaling of the sigmoid nonlinearities. This means that the benefits of training with noise can be obtained without the computational cost of averaging over many noisy samples. These results provide justification for gain scaling as a heuristic for improving generalization. Application of this technique to a single-hidden-layer perceptron with linear output is considered.>