A cluster-filter feature selection approach

Vimal Kumar Dubey, Amit Kumar Saxena, Madan Madhaw Shrivas · 2016

In this paper, a feature selection method is presented for the multiclass data sets. This method is the hybridization of k-means clustering using cosine similarity as a distance measure and information Gain. In the method unsupervised Cosine Similarity is used for grouping of features i.e. K-means clustering is used to make a cluster of features and then information gain is employed to select a most relevant feature from each cluster. The dataset with the selected feature is tested for classification accuracy with cross - validation approach. Three classifiers namely Naïve Bayes (NB), K-Nearest Neighbor and Classification and Regression trees (CART) has been used as the base classifiers for getting classification accuracy. Obtained results are compared with filter-based feature selection technique (Information Gain).

Read the paper · More papers on PaperTik