Assessing the significance of performance differences on the PASCAL VOC challenges via bootstrapping
Mark Everingham, S. M. Ali Eslami, Luc Van Gool, Christopher K. I. Williams, John Michael Winn, Andrew Zisserman · Lirias · 2013
In the PASCAL VOC challenges, entrants in a par-ticular competition are evaluated in terms of a specified metric. It can happen that some entrants will have sim-ilar scores, and it is of interest to assess the significance of these differences. For example, we might be interested to know if the highest-scoring entry is significantly bet-ter than some of the others. In this note we discuss the use of bootstrap sampling to address this question. We first came across the idea of bootstrapping precision-recall curves in the blog comment by O’Connor (2010), although bootstrapping of ROC curves has been dis-cussed by many authors, e.g. Hall et al (2004); Bertail et al (2009). In the bootstrap (see e.g. Wasserman, 2004, Ch. 8), the data points (images in our case) are sampled with replacement from the original n test points to produce B bootstrap replicates. To compare two methods A and Mark Everingham, who died in 2012, was the key member of the VOC project. His contribution was crucial and substantial. For these reasons he is included as the posthumous first author of this paper. An appreciation of his life and work can be found in Zisserman et al (2012).