Breast Cancer Classification Using Machine Learning and Feature Subset Selection Techniques

Sujata Ray, Debasmita Pradhan, Niranjan Kumar Ray · 2024

Breast cancer is one of the most common cancers among women worldwide and early detection plays a vital role to reduce the mortality rate. In this study, we propose a novel machine learning-based classification model for breast cancer classification, combining feature selection techniques and clas-sification algorithms. The Wisconsin Diagnostic Breast Cancer (WDBC) dataset is used, where feature selection is performed using Fisher Discriminant Ratio (FDR) and Pearson Correlation Coefficient (PCC). The selected features are then used to train and evaluate Support Vector Machine (SVM) and XGBoost classifiers. From the results it is observed that SVM, with features selected using fisher discriminant ratio performs better than other compared models with an accuracy of 97.66% and a recall of 98.52%, crucial for accurately identifying malignant cases. It is also observed that features like Radius mean, Perimeter mean, Compactness mean, Concative mean, Fractal dimension standard error, Texture worst value (mean of the three largest values) across all cells, Smoothness worst value (mean of the three largest values) across all cells, Compactness worst value (mean of the three largest values) across all cells are important features of breast cells for classifying the data into Malignant and Benign. This work emphasizes the potential of combining feature selection and machine learning for more accurate and efficient cancer diagnosis.

Read the paper · More papers on PaperTik