A Comparative Study for Different Resampling Techniques for Imbalanced datasets
Alaa Mahmoud Elsobky, Arabi El-Said Keshk, Mohamed G. Malhat · IJCI International Journal of Computers and Information · 2023
The imbalanced data is a significant challenge forresearchers in supervised machine learning. Current data mining algorithms are not effective for processing imbalanced data.In fact, this problem reduces classification accuracy because theprediction of minority classes is inaccurate. The classificationof imbalanced data is the major challenge that has receivedsignificant attention. Therefore, The use of sampling techniquesto improve classification performance has been a significantconsideration in related work. In this paper, a comparativestudy of six different sampling algorithms is performed. Theemployed sampling algorithms are from different samplingtechniques: two oversampling algorithms, two undersamplingalgorithms, and two combination algorithms between oversampling and undersampling. The techniques used in oversamplingare random oversampling and SMOTE, while undersamplingtechniques are random undersampling and a near miss. Acombination of oversampling and undersampling techniquesis SMOTE TOMEK and SMOTEEN. This comparative studyaims to examine the impact of the employed sampling method.Algorithms on the performance of three classifiers: SVM, KNN,and logistic regression. Cross-validation experiments on 12standard datasets show that the SMOTEEN sampling Thealgorithm achieves significant improvements compared withother typical algorithms.