N-gram Feature Selection for Text Classification Based on Symmetrical Conditional Probability and TF-IDF

Woo-sik Choi, Seoung Bum Kim · Journal of Korean Institute of Industrial Engineers · 2015

The rapid growth of the World Wide Web and online information services has generated and made accessible a huge number of text documents. To analyze texts, selecting important keywords is an essential step. In this paper, we propose a feature selection method that combines a term frequency-inverse document frequency technique and symmetrical conditional probability. The proposed method can identify features with N-gram, the sequential multiword. The effectiveness of the proposed method is demonstrated through a real text data from the machine learning repository, University of California, Irvine.

Read the paper · More papers on PaperTik