Fast Feature subset selection algorithm based on clustering for high dimensional data
S. D. Potdukhe, Zes Coer · 2014
A Feature selection algorithm employ for removing irrelevant, redundant information from the data. Amongst feature subset selection algorithm filter methods are used because of its generality and are usually good choice when numbers of features are large. In cluster analysis, graph-theoretic clustering methods to features are used. In particular, the minimum spanning tree (MST)- based clustering algorithms are adopted. A Fast clustering bAsed feature Selection algoriThm (FAST) is based on MST method. In the FAST algorithm, features are divided into clusters by using graph-theoretic clustering methods and then, the most representative feature that is strongly related to target classes is selected. Features in different clusters are relatively independent. A feature subset selection algorithm (FAST) is used to test high dimensional available image, microarray, and text data sets. Traditionally, feature subset selection research has focused on searching for relevant features. The clustering-based strategy of FAST having a high probability of producing a subset of useful and independent features.