Maximum Entropy Based Bengali Part of Speech Tagging

Asif Ekbal, Rejwanul Haque, Sivaji Bandyopadhyay · Research in Computing Science · 2008

Part of Speech (POS) tagging can be described as a task of doing automatic annotation of syntactic categories for each word in a text document. This paper presents a POS tagger for Bengali using the statistical Maximum Entropy (ME) model. The system makes use of the dierent contextual information of the words along with the variety of features that are helpful in predicting the various POS classes. The POS tagger has been trained with a training corpus of 72, 341 word forms and it uses a tagset 1 of 26 dierent POS tags, dened for the Indian languages. A part of this corpus has been selected as the development set in order to nd out the best set of features for POS tagging in Ben- gali. The POS tagger has demonstrated an accuracy of 88.2% for a test set of 20K word forms. It has been experimentally veried that the lexi- con, named entity recognizer and dierent word suxes are eective in handling the unknown word problems and improve the accuracy of the POS tagger signicantly. Performance of this system has been compared with a Hidden Markov Model (HMM) based POS tagger and it has been shown that the proposed ME based POS tagger outperforms the HMM based tagger.

Read the paper · More papers on PaperTik