On the Use of Evidence in Neural Networks
David H. Wolpert · 1992
1 THE EVIDENCE APPROXIMATION It has recently become popular to consider the problem of training neural nets from a Bayesian viewpoint (Buntine and Weigend 1991, MacKay 1992). The usual way of doing this starts by assuming that there is some underlying target function f from R n to R, parameterized by an N-dimensional weight vector w. We are provided with a training set L of noise-corrupted samples of f. Our goal is to make a guess for w, basing that guess only on L. Now assume we have i.i.d. additive gaussian noise resulting in P(L | w, b) µ exp(-b c 2 (w, L)), where c 2 (w, L) is the usual sum-squared training set error, and b reflects the noise level. Assume further that P(w | a) µ exp(-aW(w)), where W(w) is the sum of the squares of the weights. If the values of a and b are known and fixed, to the values a t and b t respectively, then P(w) = P(w | a t ) and P(L | w) = P(L | w, b t ). Bayes' theorem then The Bayesian "evidence" approximation, which is closely related to gene...