Oversampling Techniques for Diabetes Classification: a Comparative Study

Francisco Mesquita, José Maurício, Gonçalo Marques · 2021 International Conference on e-Health and Bioengineering (EHB) · 2021

This paper presents a comparative study between oversampling variants and machine learning (ML) algorithms to discover the better method for predicting diabetes using unbalanced data. “PIMA Indian Diabetes Data Set” is an unbalanced dataset used by numerous researchers to study ML applied to diabetes. This work presents a comparative analysis between different oversampling variants and the ML algorithms to predict diabetes disease. Ten ML methods and six oversampling techniques have been tested to solve the problem of unbalanced data and make a correct prediction of diabetes disease. The results recommend the combination of SVM SMOTE and AdaBoost algorithm with 83.12%, 73.53%, 86.21%, 90.70% and 26.47% Accuracy, Recall, Precision, Specificity and False Negative Rate, respectively.

Read the paper · More papers on PaperTik