Empirical Study of Feature Selection Methods for High Dimensional Data

S Deepalakshmi, Thambusamy Velmurugan · Indian Journal of Science and Technology · 2016

Background/Objectives: Feature Selection is a process of selecting features that are relevant which is used in model ­construction by removing redundant, irrelevant and noisy data. A typical application of Text Mining is classification of messages and e-mails into spam and ham. Methods/Statistical Analysis: This article gives a comprehensive overview of the various Feature Selection methods for Text Mining. Various Filter methods like Pearson Correlation, Chi-square, Symmetrical Uncertainty and Mutual Information are applied to select the optimal set of features. Findings: Filter Feature Selection methods are used to classify Text data. Various Classification algorithms are applied using the optimal set of ­features obtained. The accuracy of classification algorithms are verified based on the chosen data set. Novelty/ Improvements: A comparative study of various filter methods for Feature Selection and classification algorithms for performance evaluation is conceded in this research work.Keywords: Chi-Square, Feature Selection, Filter Method, Mutual Information, Pearson Correlation

Read the paper · More papers on PaperTik