Imbalance Dataset Handling for Classification using Machine Learning Algorithm

S. V., Nanditha Krishna, Kavita Avinash Patil, Salna Joy, R Baby Chithra, Raghavendra Patil G E · Nanotechnology Perceptions · 2024

Machine learning models are severely impacted by unbalanced data. Models may become biased in favor of the majority class when one class has a significantly larger sample count than the other, which would hurt the minority class's performance. Overfitting is made more likely by imbalanced data since the model may become accustomed to memorizing most class samples rather than identifying underlying patterns. In order to tackle these problems in the classification space, this work investigates a number of approaches, such as cost-sensitive learning, SMOTE, under- and oversampling, and ensemble deep learning techniques. They assess how well these techniques work with various datasets and offer details on their advantages and disadvantages. A taxonomy of approaches, including as resampling, algorithmic and other approaches, for imbalanced binary and multi-class classification issues is presented in the paper.

Read the paper · More papers on PaperTik