A Hybrid POS Tagger for Indian Languages

Momen Mohamed, Samar Sinha · 2011

This paper describes the work on building Part-of-Speech (POS) tagger for 12 Indian Languages using hybrid approach, and presents the performance of the tagger for each Indian language. Unlike the most of the previous POS taggers for Indian languages which are designed to annotate few languages, the present tagger called 'POS Tagger' is an attempt to facilitate annotation of several Indian languages following a computational approach. The POS Tagger is trained on 80K to 85K tagged corpora for each language from the LDC-IL corpus. Finally, this paper highlights the performance of the tagger and the need of language specific resources required for obtaining optimal result.

Read the paper · More papers on PaperTik