An AdaptiveSMOTE-RF algorithm for unbalanced data classification
Tengfei Cao, Xuechun Liang, Lingnan Xie · 2024
The SMOTE algorithm selects a neighborhood near a minor sample and ignores the distribution characteristics of a small number of samples. This marginalizes the minority samples and affects the model’s classification performance. To mitigate this, Adaptive-SMOTE and random forest feature selection are combined to reduce the impact of data imbalance during data pre-processing. In the data pre-processing stage, the adaptive and heavy-sampling SMOTE algorithm determines the sampling ratio from synthetic samples in an adaptive manner to better control the sampling process. Next, optimal features are selected through random forest feature selection methods, and each tree is assigned a corresponding weight. This empowers trees with better classification capabilities, giving them greater voting influence during the voting stage and enhancing the random forest algorithm’s overall classification performance. Experiments on the Kaggle dataset demonstrate that, compared to the original algorithm, the improved algorithm enhances the accuracy and overall sample classification performance of the minority samples.