Unbalanced data classification based on oversampling and integrated learning

Yongjun Zhang, Xiaowen Jian · 2021

A lot of unbalanced data exist in the real world, but most of traditional classification algorithms assume that the data is balanced and the misclassification cost is the same with all classes. Therefore, traditional classification algorithm should be modified. This paper proposed an improved unbalanced data classification algorithm: WSMOTEBoost which based on SMOTE and AdaBoost integrated learning. Add two weight improvements, in the level of data processing, considering the SMOTE adding boundary and the central sample. The sampling weight of fundamental sample was comprehensively determined using the Euclidean distance and the weight of iterative sample in AdaBoost, so that the samples with high value will sampled more. In terms of classification algorithm, cost-sensitive learning is introduced, sensitive factors are used to measure the classification cost, and the error rate of specific classes is used to improve the original Adaboost sample weight iteration, it made AdaBoost more concerned with the minority class that were misclassified. Experiments on three open unbalanced data sets of UCI show that the proposed method has more efficiency compared to AdaBoost and SMOTE Boost algorithms in F1 value and AUC. It is more suitable for unbalanced data classification.

Read the paper · More papers on PaperTik