Automated part-of-speech analysis of Urdu: conceptual and technical issues.
Andrew Hardie · Lancaster EPrints (Lancaster University) · 2005
Part-of-speech (POS) tagging is the process of labelling tokens in a text with tags that indicate their morphosyntactic category, and has a wide range of applications in computational and corpus linguistics, such as the production of corpus-based dictionaries and grammars.This paper describes an experiment in extending POS tagging to a hitherto untagged language, Urdu.The most challenging task in POS tagging is disambiguation, i.e. the resolution of the contextual ambiguity of a token for which more than one tag is possible.Three important approaches to disambiguation have been developed: approaches based on rules devised by a linguist; probabilistic approaches based on the application of corpus-derived statistics in a mathematical model such as a Markov model; and Brill (1995)'s approach where rules are learned automatically from a corpus.However, given that only a small amount of pre-tagged data was available for Urdu, only the rulebased approach was appropriate for the Urdu tagger described here.A rule-based tagger for Urdu was created within the Unitag architecture, together with the requisite language-specific resources for Urdu (including a tagset, an analyser, a lexicon, and a rule list).An evaluation of the tagger suggests that it performs at a level of accuracy notably below that commonly reported for languages such as English.However, this poor performance is primarily attributable to the small size of the lexicon, which is attributable to the small quantity of training data available.The rule-based disambiguation rules was more successful.