A Cluster Switching Method for Sampling Imbalanced Data

Wanthanee Prachuabsupakij, Supaporn Simcharoen · 2018

Classification on imbalanced data is one of the most interesting in data mining challenge. In this paper, a new repetitive sampling method, namely ClusIM is proposed to improve the prediction performance on imbalanced dataset using Clustering Switching Method based on K-means algorithm in order to generate new subset in attempting to reduce the overlapping between the minority class instances and majority class instances in each subset. Then, SMOTE algorithm is used to operate on each subset according to imbalance ratio of the subset. It generates the synthetic instances of the minority class. The ClusIM will generate two-balanced final training set, which are classified using SVM and to combine the model through maximum probability vote. Our experiments are based on six imbalanced data sets from UCI and one real-world dataset, comparing ClusIM with four well-known classification algorithms (SVM, Bagging, AdaboostM1, and AdaCost). Amongst the compared algorithms, ClusIM has higher F-measure and G-mean results than the other methods. This study supported that ClusIM is capable of improving the performance of learning algorithm on imbalanced dataset.

Read the paper · More papers on PaperTik