Analyzing Oversampling and Machine Learning Approaches for Imbalanced Dataset Classification
Dini Adni Navastara, Chastine Fatichah, Yulia Niza, Fiqey Indriati Eka Sari, Muchamad Maroqi Abdul Jalil · 2023
Imbalanced data, characterized by a substantial difference in data distribution between majority and minority classes, poses a critical challenge in predictive modeling. This disparity often leads to the misclassification of the minority class, which may contain vital information for real-world applications. Consequently, addressing imbalanced data is paramount, given its potential repercussions in critical classification scenarios. In this study, we conducted a comprehensive analysis of oversampling and ensemble learning techniques to mitigate imbalanced data issues. Through an extensive evaluation process employing confusion matrices, we measured the performance of these methods across various binary datasets. Remarkably, our findings showcased the efficacy of these techniques, with standout results such as a recall score of 0.6883 for the Haberman’s Survival dataset using the KNN classification method in conjunction with Borderline-SMOTE oversampling, a recall score of 0.8391 for the COVID19 dataset with the KNN classification method and SMOTE oversampling, and an impressive recall value of 0.9476 for the Credit Card Fraud dataset when applying the XGBoost classification method and ROS oversampling. It is important to note that the performance outcomes are intrinsically tied to the unique characteristics of each dataset. This study provides insights for handling imbalanced data based on dataset characteristics, aiding predictive modeling in real-world scenarios.