Breast Cancer Patient Survival Prediction Using Machine Learning on An Imbalanced Dataset

Subrata Saha, Md. Motinur Rahman, Md Shamimul Islam, Md Mahmudul Hasan, Adity Bhowmik, Mohammad Abu Sayid Haque · 2024

Breast cancer (BC) is a leading cause of cancer-related deaths globally in women. Early detection and treatment of BC can greatly lower death rates. Various methods have been suggested to help in the early identification and treatment of BC. Similarly, studies on breast cancer survival are important because predicting the survival of BC patients can aid healthcare providers in efficient resource allocation. Furthermore, this can empower patients to make informed decisions about their treatment options and plans for the future. This study explores the efficacy of Machine Learning (ML) algorithms in predicting BC patient survival, utilizing the NCI SEER breast cancer dataset. The data was preprocessed by handling missing values and normalization. Moreover, the Adaptive Synthetic (ADASYN) algorithm was implemented as the dataset was imbalanced. Then six ML models, including K Nearest Neighbors (KNN), Multilayer Perceptron (MLP), Voting classifier (VC), Extreme Gradient Boosting (XGB), Decision Tree (DT), and Adaptive Boosting (AdaBoost) were implemented and compared. Model evaluation was conducted using metrics such as Accuracy, Precision, Recall, F1-score, Confusion Matrix, and the AUC-ROC. The XGB model outperformed others by gaining the highest accuracy of 93.8 % and AUC-ROC of 0.98, highlighting its potential for a reliable survival prediction model. The unique contribution of this experiment is its approach to processing an imbalanced dataset, which leads to a substantial improvement in accuracy. The result also indicates the patient's expected survivability is highly correlated with their age, race, cancer stage, and, number of regional positive lymph nodes.

Read the paper · More papers on PaperTik