Feature Subset Selection in High Dimensional Data Using Clustering
Sudheer Babu Nadakudur, Homer Benny Bandela · 2014
Feature selection is a process of identifying a subset of potential features that can be used to produce useful results similar to the original feature set. The feature selection should be characterized and viewed viewed from the efficiency and effectiveness point of view. Efficiency is a matter of time needed to find a subset of features where as effectiveness deals with the quality of the subset of features. These two factors are studied using the fast clustering-based feature selection algorithm (FAST). According to this algorithm the features are divided into clusters first and then the most representative feature that strongly relates to the target class is selected from each cluster to form a subset of features. The FAST algorithm has the capability of generating a feature subset of potential features even though the features from different clusters are independent from each other. The efficiency and effectiveness of the FAST algorithm are evaluated through an empirical study.