An Introduction to Bayesian Statistics, Part 2: Learning from Data
Tom Fearn · NIR news · 2012
A simple example Suppose I have four identical looking opaque bags containing red and green counters. I tell you, truthfully, that the proportions of red counters in the four bags are 1 / 5, 2 / 5, 3 / 5 and 4 / 5. I then show you one of the bags and tell you, again truthfully, that this one has been chosen from the four using a physical randomisation device (e.g. two coin tosses) that gave each one equal probability. I do not tell you which one it is. At this point, without any further information, it would be reasonable to describe your state of knowledge about the selected bag by a probability distribution that assigns probability 1 / 4 to each of the four possibilities. Now suppose I allow you to carry out an experiment. Without seeing the contents of the bag, you draw a counter from it, observe its colour and replace it. The results of three such trials are red, red, green, or RRG with the obvious notation. Clearly these data ought to change the probabilities you assign to the four bags. It should raise the probabilities of the bags with more reds, since the observed outcome is more likely if the bag is one of these than if it is the one with only 20% reds. In fact we can use Bayes theorem to calculate what our new probabilities should be. First, we need to set up some notation. Let q be the unknown proportion of reds in the selected bag. Then the probability of drawing a red counter is q, and the probability of drawing a green one is 1 − q. The draws are independent, and we replaced the counter each time so the probabilities do not change. Thus the probability of observing RRG is q(1 − q). For q = 1 / 5, for example, this gives RRG a probability of 4 / 125, the first figure in the second column of Table 1. Table 1 shows the calculation of p(q | RRG), or equivalently p(q | data) using Bayes theorem in the form