ON DOCUMENT POPULATIONS AND MEASURES OF IR EFFECTIVENESS
Stephen E. Robertson · 2007
Work on the statistical validity of experimental results in retrieval tests has concentrated on treating the topics as a sample from a population, but regarding the collection of documents as fixed. This paper raises the argument that we should also consider the documents as having been sampled from a population. It follows that we should regard a per-topic measurement as also having a per-topic noise or error associated with it, which may depend critically on the number of relevant documents for that topic. Some of the common measures used in retrieval testing are re-examined from this point of view. The examination is essentially theoretical, supported by limited simulation experiments. 1