An Enhanced Feature Selection Method for Large-Scale Dataset with Numerous Similar Features
Wenting Chen, Ming Li, Xinyuan Zhao, Meng Huang, Jian Zhang, Heng Zhang · 2025
Feature selection is regarded as a vital method for dimensionality reduction, aimed at removing redundant and irrelevant features while retaining the most informative ones from the original dataset. However, large-scale datasets often contain numerous similar features and significant noise, thereby degrading performance of feature selection methods. To address this issue, we propose a novel feature selection method specifically designed for large-scale datasets, called the enhanced Relief-F (ERelief-F) algorithm. Specifically, our method mainly incorporates two key strategies: 1) by employing the hierarchical clustering algorithm, the preliminary feature selection is established so as to obtain the representative feature subset; 2) by integrating the sample similarity with the Relief-F algorithm, we develop a more robust version of Relief-F on the representative feature subset, allowing us to assign accurate weights that reflect feature importance. Experimental results on six public datasets demonstrate that the proposed ERelief-F method outperforms state-of-the-art feature selection techniques in terms of both accuracy and robustness.