CUSBoost: Cluster-Based Under-Sampling with Boosting for Imbalanced Classification
Farshid Rayhan, Sajid Ahmed, Asif Mahbub, Rafsan Jani, Swakkhar Shatabda, Dewan Md. Farid · 2017
Class imbalance classification is a demanding research problem in the context of machine learning and its applications, as most of the real-life datasets are often imbalanced in nature. Existing learning algorithms maximise the classification accuracy by correctly classifying the majority class, but misclassify the minority class. The challenge lies in that the minority class instances represent the data of greater interest than the majority class instances in real-life applications. Recently, several techniques based on sampling methods (under-sampling and oversampling of the majority and minority class respectively), cost-sensitive learning methods, and ensemble learning have been used in the literature for classifying imbalanced datasets. In this paper, we present a new clustering-based under-sampling approach with boosting (AdaBoost) algorithm, called CUSBoost, for effective imbalanced classification. The proposed algorithm provides an alternative to RUSBoost (random under-sampling with AdaBoost) and SMOTEBoost (synthetic minority oversampling with AdaBoost) algorithms. We evaluated the performance of the CUSBoost algorithm against the state-of-the-art methods based on ensemble learning like AdaBoost, RUSBoost, SMOTEBoost on 13 imbalanced binary and multi-class datasets with various imbalance ratios. The experimental results show that the CUSBoost is a promising and effective approach for dealing with highly imbalanced datasets.