A boosting based approach to handle imbalanced data

Sahar Hassanzadeh Mostafaei, Jafar Tanha, Negin Samadi, Soodabeh Imanzadeh, Nazila Razzaghi-Asl · 2022

Most real-world datasets usually contain imbalanced data. Learning from datasets where the number of samples in one class (minority) is much smaller than in another class (majority) creates biased classifiers to the majority class. The overall prediction accuracy in imbalanced datasets is higher than 90%, while this accuracy is relatively lower for minority classes. In this paper, we first propose a new technique for under-sampling based on the Peak clustering method from majority class on imbalanced datasets. We then propose a novel boosting-based algorithm for learning from imbalanced datasets, based on a combination of the proposed Peak under-sampling algorithm and over-sampling technique (SMOTE) in the boosting procedure. In the proposed algorithm misclassified examples are not given equal weights. The proposed algorithm selects useful samples from the majority class and creates synthetic samples from the minority class to indirectly change the update weights. We designed experiments using ten datasets from various domains and different evaluation metrics that show improved prediction performance on the minority class.

Read the paper · More papers on PaperTik