Ranking and selecting terms for text categorization via SVM discriminate boundary
Tien-Fang Kuo, Yasutoshi Yajima · International Journal of Intelligent Systems · 2005
The problem of natural language document categorization consists in classifying documents into predetermined categories based on their contents. Each distinct term, or word, in the documents is a feature for representing a document. In general, the number of terms may be extremely large and the dozens of redundant terms may be included, which may deteriorate the performance of classification. In this paper, an SVM based feature ranking and selecting method for text categorization is proposed. The contribution of each term for classification is calculated based on the nonlinear discriminate boundary generated by support vector machine (SVM). The results of experiments on the Reuters-21S78 dataset show that the proposed method achieves higher classification performance than existing feature selection based on LSI and x/sup 2/ statistics values.