A novel technique to solve class imbalance problem
Israt Jahan Emu, Dilshad Jahin, Subrina Akter, Muhammed J. A. Patwary, Shamima Akter · 2022 International Conference on Innovations in Science, Engineering and Technology (ICISET) · 2022
The introduction of Big Data has proclaimed the beginning of a new age of scientific advances. One of the most often encountered problems with raw data is a class imbalance, which refers to an imbalanced distribution of response variable values. Modern machine learning approaches struggle to cope with imbalanced data by concentrating only on the majority class's misclassification rate and neglecting the minority class. This study introduces a unique oversampling strategy for obtaining a balanced dataset. In this method, new samples of the minority class are increased in such a way that the misclassification rate of the majority class remains low. The reason behind this is, the data of the dominant class is authentic, but the data of the minority class is synthesized. So our goal was not to misclassify genuine data as a result of fabricated data. Using the decision tree as a base classifier, we conducted a largescale experiment using 5-fold cross-validation on three datasets from the UCI machine learning repository, with low, medium, and high-level imbalanced ratios, respectively. For validation purposes, the datasets are classified by the decision tree classifier utilizing 5-fold cross-validation. Finally, on three UCI datasets, our approach outperformed SMOTE, with the average values of the Decision tree classifier's Sensitivity and Specificity measures for each dataset.