Comparison of TnT, Max.Ent, CRF Taggers for Urdu Language
M. Humera Khanam, K. V. Madhumurthy, Md. A. Khudhus · 2013
The development of statistical taggers for Urdu language is an important milestone toward Urdu language processing. In this paper we look at the efficient methods of computational linguistics. We did our Experiments with some of the widely used POS Tagging approaches on Urdu language. Part-of-Speech (POS) Tagging is a process that attaches each word in a sentence with a suitable tag from a given Tag set. In this paper, three stateof-art probabilistic taggers i.e. TnT tagger, Maximum Entropy tagger and CRF (Conditional Random Field) taggers are applied to the Urdu language. A training corpus of 100000 tokens is used to train the models. We compare all the three taggers with same training data and finally we concluded that CRF Tagger shows the better accuracy.