Confidence Bands for ROC Curves

Sofus A. Macskassy, Foster John Provost, Michael L. Littman · 2003

In this paper we study techniques for generating and evaluat-ing confidence bands on ROC curves. ROC curve evaluation is rapidly becoming a commonly used evaluation metric in machine learning, although evaluating ROC curves has thus far been lim-ited to studying the area under the curve (AUC) or generation of one-dimensional confidence intervals by freezing one variable— the false-positive rate, or threshold on the classification scoring function. Researchers in the medical field have long been us-ing ROC curves and have many well-studied methods for ana-lyzing such curves, including generating confidence intervals as well as simultaneous confidence bands. In this paper we intro-duce these techniques to the machine learning community and show their empirical fitness on the Covertype data set—a stan-dard machine learning benchmark from the UCI repository. We show how some of these methods work remarkably well, others are too loose, and that existing machine learning methods for gen-eration of 1-dimensional confidence intervals do not translate well to generation of simultanous bands—their bands are too tight. 1.

Read the paper · More papers on PaperTik