Imbalanced Data Classification Based on a Hybrid Resampling SVM Method
Lu Cao, Yikui Zhai · 2015
Imbalanced datasets are frequently founded in many different applications, causing poor predication performances for minority class. In the paper, a hybrid re-sampling approach was proposed to deal with the two-class imbalanced data classification. Firstly, SMOTE technique is used to generate synthetic points for the minority class, then, under-sampling technique was used to delete some samples of the majority with less classified information. Thus, relative balanced training datasets are generated and we use SVM to cope with the new dataset. Experimental results on a synthetic dataset and five benchmark UCI datasets are provided to show the effectiveness of the proposed method.