Explainable Multi-Model Framework for Breast Cancer Diagnosis Using Heterogeneous Clinical Data

MT. Shourovy Akter, Md Istiaq Mohhamad Shuvo, Sabekun Nahar Tithe, Md Naimur Islam Nissan, Pankaj Bhowmik, Md. Delowar Hossain · 2025

In today's world, early detection of diseases has become more important due to the rapid growth of the global population and the corresponding medical challenges. Hence, we proposed a machine learning pipeline that incorporates detailed preprocessing, including exploratory data analysis (EDA), median imputation to handle missing values and removal of zero variance and highly correlated features. In the feature selection stage, we used ANOVA, Boruta and RFE to identify the most relevant features and unified their results. We trained seven different machine learning classifiers, namely Logistic Regression (LR), Decision Tree (DT), Support Vector Machine (SVM), K-Nearest Neighbor (KNN), XGBoost, LightGBM (LGBM), and Gaussian Naive Bayes (GNB). Performance was evaluated based on accuracy, precision, recall, F1-score, Matthews Correlation Coefficient (MCC) and Receiver Operating Characteristic (ROC) curves. The SVM model achieved the highest accuracy (98.25%) on the Wisconsin Diagnostic Breast Cancer (WDBC) dataset, while KNN (k=7) achieved the highest accuracy (99.29%) on the Wisconsin Breast Cancer Original (WBCO) dataset. In addition, SHAP summary and force plots were used to interpret the contribution of features and explain model predictions for both datasets.

Read the paper · More papers on PaperTik