A new feature selection method based on clustering

Huawen Liu, Yuchang Mo, Jiyi Wang, Jianmin Zhao · 2011

Feature selection is an effective technique to put the high dimension of data down, which is prevailing in many application domains, such as text categorization and bio-informatics, and can bring many advantages, such as improving efficiency and avoiding over-fitting, to learning algorithms. Currently, many efforts have been attempted in this field and various feature selection methods have been developed and proved to be very competitive. Unlike other selection methods, in this paper we propose a new method to select important features using a manner of feature clustering. The main character of our method is that it works like data clustering in an agglomerative way. In this method, each feature is considered as a data point clustered with between-cluster and within-cluster distances. As a result, the selected feature subset has minimal redundancy among its members and maximal relevance with the class labels. Our performance evaluations on seven benchmark datasets show that the classification performance achieved by our proposed method is better than other feature selection methods.

Read the paper · More papers on PaperTik