Feature Subset Selection Algorithm for Large Volumes of Data Based on Clustering
I P. Manimaran, Mercy Paul Selvan, K. S. Rangasamy · 2014
Clustering which tries to group a set of points into clusters such that points in the same cluster are more similar to each other than points in different type of clusters. In the generative clustering model, the form of parametric data generation is assumed, and the main goal in the maximum likelihood formulation is to find the parameters that maximize the probability (likelihood) of generation of the data given the model. The FAST algorithm works in two steps. The first step of the algorithm is, features are divided into clusters by using graph-theoretic clustering methods. The second step, the most representative feature that is strongly related to target classes is selected from each cluster to form a subset of features. The Features in the different clusters are relatively independent the clustering-based strategy of FAST has a high probability of producing a subset of useful and independent features. To ensure the efficiency of FAST, To assume the efficient minimum-spanning tree (MST) clustering method. In clustering process, semi-supervised learning is a class of machine learning techniques that make use of both labeled and unlabeled data for training - typically a small amount of labeled data with a large amount of unlabeled data. Semi-supervised learning falls between without any labeled training data and with completely labeled training data. Feature selection involves identifying a subset of the most useful features that produces compatible results as the original entire set of features.