Comparing feature sets for learning text categorization

Martijn Spitters · RIAO Conference · 2000

This paper describes an experimental study of feature selection in statistical learning of text categorization. We used the χ2-statistic to select the most distinguishing terms, term-bigrams, and term-trigrams as the text features. We found that applying syntactic restrictions to the bigrams and trigrams as an additional selection method enhances precision. Combining terms, bigrams, and trigrams into one feature set leads to a substantial improvement of the categorization results. We evaluated two machine learning algorithms for the task of text categorization: TiMBL (a memory-based classifier) and c5.0 (a decision tree learning algorithm).

Read the paper · More papers on PaperTik