Bayesian Backpropagation Over I-O Functions Rather Than Weights

David H. Wolpert · 1993

1 INTRODUCTION In the conventional Bayesian view of backpropagation (BP) (Buntine and Weigend, 1991; Nowlan and Hinton, 1994; MacKay, 1992; Wolpert, 1993), one starts with the "likelihood" conditional distribution P(training set = t | weight vector w) and the "prior" distribution P(w). As an example, in regression one might have a "Gaussian likelihood", P(t | w) µ exp[-c 2 (w, t)] º P i exp [-{net(w, t X (i)) - t y (i)} 2 / 2s 2 ] for some constant s. (t X (i) and t Y (i) are the successive input and output values in the training set respectively, and net(w, .) is the function, induced by w, taking input neuron values to output neuron values.) As another example, the "weight decay" (Gaussian) prior is P(w) µ exp(-a(w 2 )) for some constant a. Bayes' theorem tells us that P(w | t) µ P(t | w) P(w). Accordingly, the most probable weight given the data - the "maximum a posteriori" (MAP) w - is the mode over w of P(t | w) P(w), which equals the mode over w of the "cost function" ...

Read the paper · More papers on PaperTik