Implementation and Evaluation of Scalable Approaches for Automatic Chinese Text Categorization
Jyh-Jong Tsay, Jing‐Doo Wang, Chun-Fu Pai, Ming-Kuen Tsay · Institutional Repositories DataBase (IRDB) · 1999
The purpose of this research is to identify scalable approaches that can handle large amount of training data such as several years of news articles, and automatically assign predefined category to Chinese free text documents. Our approach consists of the following processes: (i) term extraction, (ii) term selection, and (iii) document classification. The approach first builds a recently developed SB-tree to identify all repeated substrings, called patterns, from the text. We then proceed to identify possible boundary of terms appearing in the identified patterns. After terms are extracted from the training articles, we run term selection algorithms to select the most significant terms and to reduce the number of terms to an acceptable level. The selected terms are med by the classifier to assign a predefined category to each text document. Our current experiment uses CNA one year news as training data, which consists of 73,420 articles and is far more than previous related research. In the experiment, we implement and compare four term selection methods, the odds ratio method, the mutual information method, the information gain method and the x 2-test method, when they are combined with the naive Bayes classifier.