POS tagger for Urdu using Stochastic approaches
Vaishali Gupta, Nisheeth Joshi, Iti Mathur · 2016
Part-of-Speech tagging is a problem of Natural language processing. It is a process of labeling an accurate part of speech for each word of a given corpus sentence. There are various approaches like rule based, stochastic and hybrid that are mainly used for automatic tagging. In this paper, we shall discuss stochastic approaches for part of speech tagging of Urdu language. To develop the Urdu POS tagger, a new draft version of Urdu POS tagset is used which is designed by TDIL. Using this tagset, the CRF and HMM based part-of-speech taggers are developed for Urdu. These tagger achieved 83.37% and 81.07% accuracy respectively.