Co-relation-based Feature Extraction to Improve Classification Accuracy
Md. Rakibul Islam, Md. Awinul Hoque Utsha, Mubin Ul Haque, Ezaz Mahmud Jim, Yeasir Ramim, Md. Mehedi Hasan Hridoy · 2024
Large datasets often contain redundant or irrelevant attributes that negatively impact classifier performance. Effective feature extraction is crucial for improving classification accuracy by identifying relevant features and reducing redundancy. Traditional techniques, such as Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA), are commonly used but face limitations in balancing dimensionality reduction with feature relevance, often leading to suboptimal results in terms of classification accuracy. To address these challenges, this paper introduces a novel hybrid feature extraction approach that combines the strengths of correlation analysis and PCA. This hybrid method aims to create a minimal redundancy maximal relevance feature subset, overcoming the limitations of traditional methods by enhancing both efficiency and effectiveness in feature extraction. The proposed approach was evaluated on five diverse classification datasets—Rice, Dry Bean, Ionosphere, Breast Cancer, and Yeast—characterized by varying instances, features, and class distributions, ensuring a robust assessment of its applicability. Experimental results demonstrate that the hybrid approach significantly outperforms traditional methods in terms of classification accuracy and feature subset relevance. These findings underscore the novel contributions of this work, which include the integration of correlation analysis with PCA to improve dimensionality reduction, the creation of more relevant feature subsets, and the enhancement of classification model performance across diverse datasets and problem domains. This study highlights the proposed method’s potential as a robust feature extraction tool, offering substantial improvements in machine learning applications.