Performance and Statistical Evaluation of Three Sampling Approaches in Handling Binary Imbalanced Data Sets
Fhira Nhita, Adiwijaya Adiwijaya, Isman Kurniawan · 2023
The issue of imbalanced data is still extensively investigated because of the existence of this issue in a wide range of cases. Several methods have been proposed to address this issue and improve classification performances, especially for the minor class. However, no single method gives a satisfying result for every case. Hence, a comparative study is required to understand better the correlation between the imbalanced data characteristic and a suitable method. In this study, we statistically evaluated 21 sampling methods on 100 imbalanced data sets with a range of imbalance ratio (IR) values varying from 1.82 to 129.44. Those methods were selected to represent three sampling approaches, i.e., undersampling, oversampling, and hybrid sampling. According to the results, we found that the implementation of sampling methods significantly impacted the classification result. Also, three methods, i.e., RandomUnderSampler (RUS), InstanceHardnessThreshold (IHT), and RandomOverSampler (ROS), give the best performance with the value of balanced accuracy score (BAS) are 0.8599, 0.8460, and 0.8049, respectively. The statistical evaluation shows a significant performance difference between RUS and ROS. Meanwhile, there is no significant difference in performance between RUS and IHT, and ROS and IHT. Furthermore, the results indicated that the imbalance rate and instances per attribute ratio alone could not be used to determine the complexity level of the data set.