BRILL'S TRANSFORMATION RULE-BASED TAGGER FOR SANSKRIT

R.J. Rama Sree · Research Journal of Information Technology & Software Management · 2014

ABSTRACT - It may be noted that the growth of NLP applications in Sanskrit is hindered due to unavailability of online lexical resources. It is very difficult to prepare online lexical resources as it is a time consuming task and needs lot of expertise in both computer science and Sanskrit Grammar. One such lexical resource is annotated corpus for Sanskrit. Hence building a large annotated corpora which is very much essential for advanced NLP processing tasks and it is a prime step towards building NLP applications with less time and cost. One goal of this work is to build a data-driven POS tagger for Sanskrit using existing algorithms which are language and tag set independent. It is very easy to apply these taggers to new languages, domains, tag sets provided we have correctly annotated training corpus. They are freely available on the net. One such tagger is Brill's tagger. The accuracy of the tagger when applied to different European languages is above 95%. In this paper Brill's Transformation Rule-based Learning (TBL) tagger was tested and adapted for Sanskrit. It was shown that the present system does not obtain a very high accuracy but results were still promising with an average accuracy of 60%. Key Words : Sanskrit POS tagging, Annotating Sanskrit texts, Transformation Based Learning

Read the paper · More papers on PaperTik