TURKISH TEXT CATEGORIZATION USING N-GRAM WORD

Aysun Güran, Selim Akyokuş, Nilgün Güler Bayazıt · 2009

An N-gram is a representation method that consists of a sequence of N-contiguous characters or words. There have been so many studies which use N-gram based representations for the traditional text classification tasks. In contrast to other languages, the studies in Turkish are limited. In this paper, we analyze text classification algorithms on a Turkish dataset by using N-gram words. We have compared several classifiers (Bayesian probabilistic classifiers, nearest neighbor classifiers and decision trees) using different types of features. We applied the classifiers on different data sets that are represented with unigram, bigram and trigram words. In the experiments, a total of 600 text documents that are assigned to six categories were used and the best success rate of 95.83% was achieved by using unigrams.

Read the paper · More papers on PaperTik