Improving Diabetes Mellitus Prediction with MICE and SMOTE for Imbalanced Data
Mohammad Abdullah, Yap Bee Wah · 2022
Diabetes Mellitus is a chronic disease, and it affects our body nerves, eyes, kidney, and heart if not treated from the early stage. Clinical data faces issue of missing values due to lost in follow up of the patients imbalanced data. The aim of this study is to develop a Diabetes Mellitus prediction model using machine learning (ML) classifiers for an imbalanced dataset missing value. The PIMA Indian Diabetes Dataset was obtained from UCI machine learning repository. Data pre-processing involves data imputation using Multivariate Imputation Chained Equation (MICE) by Permutation Mean Matching method. Then the data was balanced using SMOTE and ROSE method. Results showed that the F-measure of the ML classifiers is higher under SMOTE oversampling technique. The random forest classifier has the highest F-measure (0.8307). The important features identified based on random forest classifier were glucose, BMI, and age. SMOTE technique is a good oversampling technique for imbalanced data.