Obtaining Calibrated Probability Estimates from Support Vector Machines
Joseph Drish · 1998
In many supervised learning tasks a learned classifier automatically induces a ranking of test examples, making it possible to determine the relative likelihood that a given test example belongs to a certain class. However, for many applications this ranking is not sufficient, particularly when the classification decision is cost-sensitive. In this case, it is necessary to convert the outputs of the classifier into well-calibrated posterior probabilities. A recent paper that addresses this problem is [7], which introduces new methods for estimating the probabilities from naive Bayes and decision tree classifiers. The goal of this project is to replicate that work using Support Vector Machines (SVMs). Based on the theory of Structural Risk Minimization [5], SVMs learn a decision boundary between two classes by mapping the training examples onto a higher dimensional space and then determining the optimal separating hyperplane between that space. Given a test example x, the SVM outputs a score that provides the distance of x from the separating hyperplane. The sign of the score indicates to which class j example x belongs, where j ∈ {1,−1}. The problem of interest is how to calibrate that score into an accurate class conditional posterior probability, or P ( j|x).