Lexicon Development and POS Tagging Using a Tagged Bengali News Corpus.

Asif Ekbal, Sivaji Bandyopadhyay · 2007

Lexicon development and Part of Speech (POS) tagging are very important for almost all Natural Language Process-ing(NLP) application areas. The rapid development of these resources and tools using machine learning techniques for less computerized languages requires appropriately tagged corpus. A tagged Bengali news corpus has been developed from the web archive of a widely read Bengali newspaper. This corpus is then used for lexicon development and POS tagging. Tagged Bengali News Corpus Development Newspaper is a huge source of readily available documents. A tagged corpus has been developed from the web archive of a very well known and widely read Bengali News Pa-per. The development of the tagged Bengali news corpus includes language resource acquisition using a web crawler, language resource creation which includes HTML file clean-ing and code conversion, as well as language resource anno-tation that involves defining a tag set and subsequent tagging of the news corpus. Code conversion is necessary to convert the dynamic fonts used in the newspaper into the standard Indian Standard Code for Information Interchange (ISCII) form, which can be processed for various text processing tasks. At present, the corpus contains 34 million wordforms and it is available in both ISCII and UTF-8 formats. A news corpus, whether in Bengali or in any other lan-guage has different parts like title, date, reporter, location, body etc. To identify these parts in a news corpus, the fol-lowing tagset has been defined: header (Header of the news document), title (Headline of the news document), t1 (1st headline of the title), t2 (2nd headline of the title), date (Date of the news document), bd (Bengali date), day (Day), ed (English date), reporter (Reporter-name), agency (Agency providing news), location (the news location), body (Body of the news document), p (Paragraph), table (information in tabular form), tc (Table Column), and tr (Table row).

Read the paper · More papers on PaperTik