Dual Approach to Handling Imbalanced Class in Datasets Using Oversampling and Ensemble Learning Techniques
Yoga Pristyanto, Anggit Ferdita Nugraha, Irfan Pratama, Akhmad Dahlan, Lucky Adhikrisna Wirasakti · 2021
In the field of machine learning, the existence of class imbalances in the dataset will make the resulting model have less than optimal performance. Theoretically, the single classifier has a weakness for class imbalance conditions in the datasets because of the majority of single classifiers tend to work by recognizing patterns in the majority class the datasets are not balanced. So, the performance cannot be maximized. In this study, two approaches were introduced to deal with class imbalance conditions in the dataset. The first approach uses ADASYN as resampling while the second approach uses the Stacking algorithm as meta-learning. After conducting a test using 5 datasets with different imbalanced ratios, it shows that the proposed method produced the highest g-mean and AUC score compared to the other classification algorithms. The proposed method in this study is the stacking algorithm between the SVM and Random Forest algorithms and the addition of ADASYN in the resampling process. Hence, the proposed method can be a solution for handling class imbalance in datasets. However, this study has limitations such as the dataset used is a dataset with a binary class category. For this reason, for the future work, testing will be suggested using the imbalanced class dataset with the multiclass datasets.