Bayesian Model Averaging is not Model Combination

Tom Minka · 2000

In a recent paper, Domingos (2000) compares Bayesian model averaging (BMA) to other model combination methods on some benchmark data sets, is surprised that BMA performs worst, and suggests that BMA may be flawed. These results are actually not surprising, especially in light of an earlier paper by Domingos (1997) where it was shown that model combination works by enriching the space of hypotheses, not by approximating a Bayesian model average. And the only flaw with BMA is the belief that it is an algorithm for model combination, when it is not. Bayesian model averaging is best thought of as a method for ‘soft model selection. ’ It answers the question: “Given that all of the data so far was generated by exactly one of the hypotheses, what is the probability of observing the new pair (c,x)? ” The soft weights in BMA only reflect a statistical inability to distinguish the hypothesis based on limited data. As more data arrives, the hypotheses become more distinguishable and BMA will always focus its weight on the most probable hypothesis, just as the posterior for the mean of a Gaussian focuses ever more narrowly on the sample mean. Mathematically, we can write the BMA rule as p((c,x)|D) ∝ ∑ p((c,x),D|h)p(h) (1)

Read the paper · More papers on PaperTik