A Hybrid Approach to Classify the Multiclass Imbalanced Datasets
S. S. Sridhar, Anusuya · 2023
The imbalanced nature of a dataset often leads classifiers to show more favor towards recognizing the class with a higher number of data instances. Handling this imbalance in a multiclass environment presents an additional challenge compared to binary class environments. Previous research primarily focused on addressing the imbalance in binary class scenarios, but there is a growing shift towards addressing the imbalance in multiclass environments. In imbalanced environments, data balancing methods are used to address the imbalance issue by equalizing the number of data instances in both the major and minor class groups. This approach aims to avoid bias in learning by modifying the class distribution. Algorithmic-level methods enable the learning model to adjust its learning process towards the class with a lower number of data instances, aiming to improve the prediction rate. The existing multiclass imbalance handling methods either attempted to solve the imbalance problem either by transforming the problem into binary class problems or by approaching the problem as a whole and finds the solution. In both the cases the classification model suffers from the issues like overlapping problem, least recognition of minority class samples, ignoring the minority class samples as outlier and over fitting. This research study proposes a hybrid approach that combines both data level balancing and algorithm level tuning methods. The proposed hybrid approach combines the effort of algorithm level tuning to create a cost associated dataset which avoids the models' biasness towards majority class samples and the data level balancing method to improve the count of the least class samples count and thereby avoiding the ignorance problem and makes the learning model to have pay more attention towards the minority class samples.