N-gram based text classification for Persian newspaper corpus

Mojgan Farhoodi, Alireza Yari, Ali Sayah · 2011

Statistical n-gram language modeling is applied in many domains like speech recognition, language identification, machine translation, character recognition and topic classification. Most language modeling approaches work on n-grams of words. In this paper, we employ language models classifier based on word level n-grams for Persian text classification. The presented approach computes the occurrence probability on word sequence in training data. Then by extracting the word sequence in test data, it can predict the highest probability for related class to given news text. We show that statistical language modeling can significantly cause high classification performance. The experimental results on Hamshahri corpus show satisfactory results and n-grams of length 3 are the most useful for Persian text classification.

Read the paper · More papers on PaperTik