A pre-hoc SMOTE variants approach for mitigating dataset bias in medical ensemble learning models
Djalila Boughareb, Hazem Bensalah, Hamid Séridi · Global Journal of Computer Sciences Theory and Research · 2024
Bias in datasets, particularly with respect to gender and race, presents a critical challenge in medical machine learning models. Such biases undermine predictive accuracy and compromise fairness in clinical decision making. This study addresses gender related bias and class imbalance through the application of synthetic oversampling techniques. Several variants of the Synthetic Minority Oversampling Technique were implemented, including SMOTE, Borderline SMOTE, Support Vector Machine SMOTE, and KMeans SMOTE. These methods were applied to improve the representativeness of the dataset and enhance model fairness. An ensemble learning classifier based on Gradient Boosting Machine was developed and evaluated. Performance and fairness were assessed using multiple metrics including accuracy, recall, precision, F1 score, positive predictive value, equal opportunity difference, and disparate impact. Among the tested approaches, Borderline SMOTE demonstrated the most promising outcomes, improving predictive performance and achieving more balanced positive predictive values across gender groups. It also contributed to a meaningful reduction in fairness-related disparities. These findings suggest that targeted oversampling strategies can significantly reduce bias in predictive modeling and support more equitable clinical outcomes, reinforcing the importance of fairness-aware approaches in the development of medical artificial intelligence systems. Keywords: Bias reduction; fairness; Healthcare; machine learning; SMOTE