Multilingual text categorization using Character N-gram
Makoto Suzuki, Naohide Yamagishi, Yi-Ching Tsai, Shigeichi Hirasawa · 2008
In our previous paper, we proposed a new classification technique called the Frequency Ratio Accumulation Method (FRAM). This is a simple technique that adds up the ratios of term frequency among categories. However, in FRAM, the use of feature terms is unlimited. In the present paper, we adopt character N-gram as feature terms improving the above-described particularity of FRAM. That is to say, the proposed method is language-independent because it does not depend on the low of grammar by using character N-gram. Therefore, we can classify multi-language into some categories using only one program. Next, the proposed method is evaluated by performing several experiments. In these experiments, we classify newspaper articles from English Reuters-21578, Japanese CD-Mainichi 2002 and Chinese China Times 2005 using FRAM. As a result, we show that it has the good classification accuracy. Specifically, the recall of the proposed method is 87.8% for English, 86.0% for Japanese and 72.8% for Chinese. Although it turned out that Chinese classification accuracy was extremely low in the present experiments compared with English and Japanese, the proposed method is language-independent and provides a new perspective and has excellent potential.