Improved information gain feature selection method for Chinese text classification based on word embedding
Lei Zhu, Guijun Wang, Xianchun Zou · 2017
Feature selection is a very important part of text categorization, which can reduce the dimensionality of the text representation by vector space model (VSM), and can avoid the problem of "curse of dimensionality". Information gain (IG) feature selection algorithm is one of the most effective feature selection algorithms, but it is easy to filter out the characteristic words which have a low IG score but have a strong ability of text type identification. Meanwhile, these words are often very similar to the words of high IG score. Aiming at this defect, we propose an improved feature selection method which uses word embedding to calculate the most similar words to the current dictionary selected by IG algorithm and expand the dictionary with these words under certain regulations. Finally we achieve good experimental results in Sogou Chinese text classification corpus and Fudan Chinese text classification corpus.