Label confidence based AdaBoost algorithm
Zhe Luo, Xin Dang, Yixin Chen · 2017
AdaBoost is a well-known simple and effective boosting algorithm for classification. It, however, suffers from the overfitting problem in the case of overlapping class distributions and is very sensitive to label noise. To tackle both problems simultaneously, we consider the conditional risk as the modified loss function. This modification leads to two advantages: it is able to directly take into account label uncertainty with an associated label confidence; it introduces a “trustworthiness” measure on training samples via the Bayesian risk rule, hence the resulting classifier tends to have superior finite sample performance than the original AdaBoost when there is a large overlap between class conditional distributions. We scrutinize the re-weighting procedure and classifier combination rule to show its adaptive ability on the mitigation of class noise and overfitting. Extensive experimental results using synthetic data and real-world data sets from UCI machine learning repository are provided. The empirical study shows high competitiveness of the proposed method in predication accuracy and robustness when compared with the original AdaBoost and several existing robust AdaBoost algorithms.