Learning Optimal Threshold on Resampling Data to Deal with Class Imbalance
Nguyen Thai-Nghe, Lars Schmidt-Thieme · 2010
Class imbalance is one of the challenging problems for machine learning algorithms. When learning from highly imbalanced data, most classifiers are overwhelmed by the majority class examples, thus, their performance usually degrades. Many papers have been introduced to tackle this problem including methods for pre-processing, internal classifier processing, and post-processing – which mainly relies on the posterior probabilities. Bayesian Network (BN) is known as a classifier which can produce good posterior probabilities. In this study, we propose methods to combine resampling techniques with learning optimal threshold in BN to deal with imbalanced data. Concretely, we first rebalance the datasets by using resampling methods. We then learn the optimal threshold on the posterior probabilities produced by some Bayesian classifiers such as general BN, TAN, BAN, and Markov Blanket structure on those data. We optimize the threshold for each classifier on the holdout-set to maximize the F1-Measure, then using this threshold for final classification. We also make these classifiers become cost-sensitive by injecting the unequal misclassification costs to the threshold. Our experimental results show that the proposed methods significantly outperform the baseline Naive Bayes. These methods also perform as good as the state-of-thearts and significantly better in certain cases.