Handling class imbalance data in business domain
Md Shajalal, Mohammad Zoynul Abedin, Mohammed Mohi Uddin · 2021
The major reason is that the number of labeled data points for one specific class is extremely higher than the number of labeled data points for the other class/classes. As the classical algorithms perform with standard accuracy in the balanced dataset, randomly duplicating the minority class examples or removing random majority class data points might be one solution. Previous studies also suggested that classical algorithms could provide high accuracy on a balanced dataset after applying random sampling. The common practice is to exploit the combination of random oversampling and undersampling before applying the classification algorithm. The difference between the feature vectors of the considered sample and its nearest neighbor is exploited to create synthetic minority class examples for selection. Most classification algorithms try to learn the borderline examples of each class. It is also common for any method that classifies the borderline data points as the wrong class.