Unsupervised Feature Selection via Adaptive Feature Clustering for High-dimensional Data
Hong Bo Jia, Yubin Weng · 2022 5th International Conference on Data Science and Information Technology (DSIT) · 2022
The purpose of feature selection is to obtain a feature subset from the original data to realize dimension reduction, and at the same time, improve the performance of following machine learning task. Due to the lack of label information, unsupervised feature selection is more challenging. In this paper, the idea of sample clustering is applied to feature space. Clustering the features makes the similarity between features in the same cluster high and that from different clusters low. Subsequently, representative features are selected from each cluster to form the final feature subset. Along this way, an unsupervised feature selection algorithm which considers both the similarity and relevance of different features is proposed. Experiments in comparison with both traditional and novel algorithms on benchmark data sets are performed. The results show the superiority of the proposed method.