Rejoinder Hedging Predictions in Machine Learning

Alex Gammerman, Vladimir Vovk · The Computer Journal · 2007

We are very grateful to all discussants for their interest in our article and their comments. We will organize our response by major topics raised by them. References to our article's bibliography will be given as e.g. [A3], and references to the discussion's bibliography as e.g. [D26]. All formula numbers, such as (3), refer to the formulas in the article. As we say in the article, the two most important properties expected from confidence predictors are validity (they must tell the truth) and efficiency (the truth must be as informative as possible). Conformal predictors are automatically valid, so there is little to discuss here, but so far achieving efficiency has been an art, to a large degree, and Alexey Chervonenkis, Phil Long and Sally McClean comment on this aspect of conformal prediction. Indeed, as Prof. Chervonenkis notices, the article does not contain any theoretical results about efficiency. Such a result appears as Theorem 3.1 in our book [A3]. We use a non-conformity measure based on the nearest neighbours procedure to obtain a conformal predictor whose efficiency asymptotically approaches that of the Bayes-optimal confidence predictor. (Remember that the Bayes-optimal confidence predictor is optimized under the true probability distribution, which is unknown to the Predictor.) This result only applies to the case of classification, and it is asymptotic. Nevertheless, it is our only step towards a ‘more principled way of designing good measures of strangeness’, as Prof. McClean puts it. Her question suggests the desirability of such more principled ways; we agree and would very much welcome further results in this direction.

Read the paper · More papers on PaperTik