Direct Optimization of Margins Improves Generalization in Combined Classifiers
Llew Mason, Peter L. Bartlett, Jonathan Baxter · 1998
0 0 1 Cumulative training margin distributions for AdaBoost versus our "Direct Optimization Of Margins" (DOOM) algorithm. The dark curve is AdaBoost, the light curve is DOOM. DOOM sacrifices significant training error for improved test error (horizontal marks on margin= 0 line). 1 Introduction Many learning algorithms for pattern classification minimize some cost function of the training data, with the aim of minimizing error (the probability of misclassifying an example). One example of such a cost function is simply the classifier's error on the training data. Recent results have examined alternative cost functions that provide better error estimates in some cases. For example, results in [Bar98] show that the error of a sigmoid network classifier f(\\Delta) is no more than the sample average of the cost function sgn(` \\Gamma yf(x)) (which takes value 1 when yf(x) is no more than ` and 0 otherwise) plus a complexity penalty term that scales as kwk 1 =`, where (x; y) 2 X \\Theta f\\Si...