A systematic review on techniques of feature selection and classification for text mining
M. Sridharan, P. Sivakumar · International Journal of Business Information Systems · 2018
Nowadays, there is a quick development in the use of internet. The large amount of structured, unstructured and semi-structured forms like videos, images, audio or texts, can be shared and used on the internet by users. The main analysis of text mining is as follows: pre-processing, feature dimension reduction (feature selection or feature extraction) and text classification, clustering on the final features. In this paper, pre-processing is a step, context sensitive stemmer used to remove the stop words, different suffixes by means to reduce the words count. The unsupervised and supervised feature selection methods like document frequency, term strength, chi-square and information gain are compared to produce the best method for the web document feature selection. The classification techniques like latent semantic analysis, genetic algorithm, Rocchio's algorithm and neural networks are also compared with systematic reviews.