Comparison of Methods to Tackle Class Imbalance in Binary Classification for IoT Applications
Watcharakorn Buttijak, Krissada Suchatpattmakul, Surakarn Suksirisophak, Watcharapan Suwansantisuk · 2020
Binary classification has applications in several areas, including the Internet of Things (IoT), and can be achieved by learning a predictive model from a labelled training dataset. A training dataset that is imbalanced leads undesirably to an inaccurate classifier. In this paper, we compare performance of existing methods that tackle class imbalance. These methods are undersampling, oversampling, hybrid, SMOTE, ADASYN, and a method that does not modify the datasets. We choose support vector machine (SVM) as a classifier and estimate classification accuracy, sensitivity, and specificity under each method. The datasets in the comparison cover a wide range of applications and class ratios. The research results indicate that undersampling does not achieve a highest accuracy, sensitivity, or specificity in any benchmark dataset, and is a not a recommended method in general. A method that does not modify the datasets has the most consistent performance; it is statistically most accurate, sensitive, and specific in three, one, and five datasets, respectively, at the significance level of the test at 10%. From the comparison, we find that a suitable method to tackle class imbalance depends on a dataset at hand and that no single method universally performs well across the benchmark datasets. The research results help researchers choose a method to mitigate a detrimental effect of class imbalance, improving classification performance.