The Effects of Class Label Noise on Highly-Imbalanced Big Data

Robert K. L. Kennedy, Justin M. Johnson, Taghi M. Khoshgoftaar · 2021 IEEE 33rd International Conference on Tools with Artificial Intelligence (ICTAI) · 2021

This study explores the effects of class label noise on a highly-imbalanced big data set by injecting varying levels of class noise into a Medicare Part B fraud detection data set. Noise parameters are used to vary the total level of class noise and the proportion of class noise between the majority and minority classes. This allows us to better understand how class noise affects imbalanced data sets and where data cleaning efforts should be focused. Four popular machine learning algorithms are evaluated using six rounds of five-fold cross-validation to determine which learners are most robust to class noise. Area under the precision-recall curve (AUPRC) results shows that negative class noise, i.e. when positive instances are incorrectly labeled as negative, has the greatest adverse effect on classification performance. Statistical results show that the XGBoost learner performs significantly better than Random Forest, Multilayer Perceptron, and Logistic Regression Learners. True positive rates and true negative rates show that there is a trade-off that occurs as the noise proportion switches between the majority and negative classes. Finally, we show that the effects of class label noise can be combatted by regularizing the XGBoost learner through shallower decision trees.

Read the paper · More papers on PaperTik