Enhancing Predictive Accuracy in Medical Data Through Oversampling and Interpolation Techniques
Alma Rocío Sagaceta-Mejía, Pedro Pablo González-Pérez, Julián Fresán-Figueroa, Máximo Eduardo Sánchez-Gutiérrez · Mathematics · 2025
Class imbalance is a major challenge in supervised classification, often leading to biased predictions and limited generalization. This issue is particularly pronounced in medical diagnostics, where datasets typically contain far more negative than positive cases. In this study, we compare two oversampling strategies: the Synthetic Minority Oversampling Technique (SMOTE) and the Conditional Tabular Generative Adversarial Network (ctGAN). Using the benchmark Pima Indians Diabetes dataset, we generated balanced datasets through both methods and trained a multilayer perceptron classifier. Performance was evaluated with accuracy, precision, sensitivity, and F1 Score. The results show that both SMOTE and ctGAN improve classification on imbalanced data, with SMOTE consistently achieving superior sensitivity and F1 Score. These findings highlight the importance of selecting appropriate augmentation strategies to enhance the reliability and clinical usefulness of machine learning models in medical diagnostics.