Analyzing the Impact of Feature Correlation on Classification Acuracy of Machine Learning Model

Satabdi Mishra, Ranjan K. Pradhan · 2023

Machine Learning has been widely used in building classification models for early detection of diseases using electronic health records of patients or variety of biosensor data on the characteristics of tissues or bio-fluids. For example, machine learning algorithms have been extensively applied for cancer detection and heat disease classification using open-source clinical data repository. However, there has been increasing number of challenges and complexities in building accurate classification models from diverse healthcare datasets available today that are basically measured under a variety of clinical settings and measurement techniques. Models that were shown to attain highest accuracy often suffer from overfitting problem. A key element in enhancing the accuracy of machine learning models is selecting correlated features or eliminating redundant features associated with a given dataset. Therefore, this study examines how the degree of correlation among features affects the predictive accuracy of three popular classifier models of breast cancer detection using clinical, demographic, anthropometric and blood metabolic data (with 10 features), and cytological characteristics of breast tissues (with 11 features) of patients. It was demonstrated that eliminating highly correlated features from the input dataset does not have any negative impact on the accuracy of classification models. SVM was found to be the best classifier with highest accuracy for both the datasets. The present results suggest that increasing the number of features increases the performance of breast cancer classifiers, and there is a modest change in prediction accuracy with the removal of highly correlated features.

Read the paper · More papers on PaperTik