Comparison of Feature Reduction Methods and Machine Learning Models for Breast Cancer Diagnosis

Todor Kirilov Avramov, Dong Si · 2017

This paper builds on previously published studies that successfully apply machine learning techniques to diagnose breast cancer from image-processed nuclear features. The study compares five different classification models (Logistic Regression, Decision Tree, k-Nearest Neighbor, Linear Support Vector Machine and Cubic Support Vector Machine) and how the choice of the features used to construct these models affects their performance. The five classifiers were trained with reduced feature subsets identified by Principal Component Analysis (PCA), correlation selection, selection based on t-test significance, and Random Feature Selection. The Random Feature Selection was the most effective feature reduction technique that identified subset of features that, when used to construct the models, yielded the highest cross validation prediction accuracies. The logistic regression model trained with the 6 features identified by random selection had 97.77% accuracy, improvement of 1.5% over the logistic regression model with 3 features reported in the original papers. The linear SVM model trained with 10 features identified by random selection had 97.87% cross validated accuracy, an improvement of 0.4% over the linear model with 3 features reported in the original paper. The cubic SVM model with 11 features had similar accuracy of 97.98%. Stacking the logistic, linear SVM and cubic SVM models in an ensemble learner improved the classification accuracy to 98.56%.

Read the paper · More papers on PaperTik