Web-Based Bengali News Corpus for Lexicon Development and POS Tagging

Asif Ekbal, Sivaji Bandyopadhyay · Polibits · 2008

"Lexicon development and Part of Speech (POS)tagging are very important for almost all Natural LanguageProcessing (NLP) applications. The rapid development of theseresources and tools using machine learning techniques for lesscomputerized languages requires appropriately tagged corpus.We have used a Bengali news corpus, developed from the webarchive of a widely read Bengali newspaper. The corpus containsapproximately 34 million wordforms. This corpus is used forlexicon development without employing extensive knowledge ofthe language. We have developed the POS taggers using HiddenMarkov Model (HMM) and Support Vector Machine (SVM). Thelexicon contains around 128 thousand entries and a manual checkyields the accuracy of 79.6%. Initially, the POS taggers have beendeveloped for Bengali and shown the accuracies of 85.56%, and91.23% for HMM, and SVM, respectively. Based on the Bengalinews corpus, we identify various word-level orthographic featuresto use in the POS taggers. The lexicon and a Named EntityRecognition (NER) system, developed using this corpus, are alsoused in POS tagging. The POS taggers are then evaluated withHindi and Telugu data. Evaluation results demonstrates the factthat SVM performs better than HMM for all the three Indianlanguages."

Read the paper · More papers on PaperTik