A Hybrid Feature Selection and Stacked Generalization Model to Detect Breast Cancer
Avijit Kumar Chaudhuri, Sulekha Das, Arkadip Ray · 2023
Breast cancer is a crucial disease among all types of cancers, accounting for almost 15% of all cancer mortalities. To reduce such a vast mortality rate, early detection of the disease is essential. A fast, accurate, and interpretable machine learning (ML) model is the research subject. Fewer features reduce the computational effort and improve interpretation. A 3-phase hybrid filter wrapper feature selection approach and a stacked classification model are evaluated on the Wisconsin breast cancer dataset with 30 features and one outcome variable. Phase 1 uses a Greedy step-wise search and selects 11 features well correlated with the class but not among themselves. Phase 2 utilizes a best-first search and logistic regression learning algorithm to select six features. In Phase 3, logistic regression (LR), Naïve-Bayes (NB), Hoeffding tree (HT), support vector machine (SVM) with the polynomial kernel, multilayer perceptron network (MLN), and the stacked model are used with the six features to identify patients with or without cancer. The stacked model uses LR, NB, and HT as base classifiers and multilayer perceptron (MLP) as the meta-classifier. Data splitting, several metrics, and statistical tests are used, along with 10-fold cross-validation, to do a comparative analysis. LR, NB, and HT demonstrate improvement across performance measures on reducing the features to six. In the 50–50 split, SVM with 30 features and MLP and LR with six features record higher than 98% accuracy. The stacked model records higher than 98% accuracy with six features and 50–50 and 80–20 splits, these values and feature reduction improves upon previous studies.