A Clustering-Based Enhanced Classifi cation Algorithm for Imbalanced Data

HU Xiao-shen · 2014

Imbalanced data exist widely in the real world and their classifi cation is a hot topic in the fi eld of machine learning. A clustering-based enhanced AdaBoost algorithm was proposed to improve the poor classifi cation performance produced by the traditional algorithm in classifying the minority class of imbalanced datasets. The algorithm firstly constructs balanced training sets by the clustering-based undersampling, using K-means clustering to cluster the majority class and extract cluster centroids and then merge with all minority class instances to generate a new balanced training set. To avoid the declining of the classifi cation accuracy caused by the shortage of training sets owing to too few minority class samples, SMOTE(Synthetic Minority Oversampling Technique) combining the clustering-based undersampling was used. Next, the misclassifi cation loss function in the basic classifi er of the AdaBoost algorithm was modifi ed based on the costsensitive learning theory to assign asymmetric misclassifi cation losses to samples of different classes. The experimental results show that, the proposed algorithm makes the model training samples more representative and greatly increases the classifi cation accuracy of the minority class, keeping the overall classifi cation performance.

Read the paper · More papers on PaperTik