A Factor Based Multiple Imputation Approach to Handle Class Imbalance

Pranita Baro, Malaya Dutta Borah · Procedia Computer Science · 2023

Class imbalance and incompleteness are the two most serious problems faced in data science and machine learning when working on real-life datasets. Both of these cases have severe implications on the ability of classification algorithms to make accurate predictions. When a dataset used for training classifiers is both imbalanced as well as incomplete, the traditional approach is to address the missing data first and then handle class imbalance but it could lead to some issues such as overfitting as well as amplification of some errors due to random duplication. In this paper, an alternate factor-based multiple imputation oversampling method (FB-MIO) is proposed to handle class imbalance as well as missing values in the training dataset at the same time. First, a new factor is presented to evaluate the density of missing values belonging to the majority class with respect to the minority class in a particular region. With the help of this factor, an oscillator is developed to guide how imputation based oversampling should be carried out. Then the training set is divided into multiple smaller subsets and used the oscillator to determine whether missing values for the majority class belonging to that subsets should be imputed or not. This would help in preventing exaggerated duplication when not needed. Experiments were carried out on 27 imbalanced datasets after random addition of missing value and the F1 and AUROC scores of FB-MIO were compared to other dataset level resampling methods such as SMOTE, ADASYN, B-SMOTE etc. The effectiveness of the proposed method has been validated after experiments on benchmark datasets and the comparative results are presented in the form of average rank and number of wins.

Read the paper · More papers on PaperTik