POS tagger for Urdu using Stochastic approaches

Vaishali Gupta, Nisheeth Joshi, Iti Mathur · 2016

Part-of-Speech tagging is a problem of Natural language processing. It is a process of labeling an accurate part of speech for each word of a given corpus sentence. There are various approaches like rule based, stochastic and hybrid that are mainly used for automatic tagging. In this paper, we shall discuss stochastic approaches for part of speech tagging of Urdu language. To develop the Urdu POS tagger, a new draft version of Urdu POS tagset is used which is designed by TDIL. Using this tagset, the CRF and HMM based part-of-speech taggers are developed for Urdu. These tagger achieved 83.37% and 81.07% accuracy respectively.

Read the paper · More papers on PaperTik