SMOTE on Numeric Breast Cancer Dataset to Overcome Imbalance Class

Fiddin Yusfida A’la, Nurul Firdaus, Hartatik Hartatik, Muhammad Asri Safi’ie · 2023

The Coimbra breast cancer dataset contains a relatively small proportion of the imbalance dataset. However, it can pose challenges when developing machine learning models because the model may be biased for the majority class and have difficulty precisely predicting the minority class. This research tries to mitigate the imbalance class using SMOTE. After implementing SMOTE, we use the Random Forest algorithm to build a machine-learning model based on 10-fold cross-validation and evaluate the model using accuracy, precision, and recall. The result shows that the model accuracy increases from 76.72% to 80.47%, the precision score rises from 76.60% to 80.00%, and the recall score from 69.23% to 81.25%. Significant implications exist for enhancing the diagnosis and treatment of breast cancer by providing more accurate prognoses of patient outcomes based on their clinical and demographic characteristics.

Read the paper · More papers on PaperTik