Part-Of-Speech Tagging for Gujarati Using Conditional Random Fields
Chirag Indravadanbhai Patel, Karthik Gali · International Joint Conference on Natural Language Processing · 2008
This paper describes a machine learning algorithm for Gujarati Part of Speech Tagging. The machine learning part is performed using a CRF model. The features given to CRF are properly chosen keeping the linguistic aspect of Gujarati in mind. As Gujarati is currently a less privileged language in the sense of being resource poor, manually tagged data is only around 600 sentences. The tagset contains 26 different tags which is the standard Indian Language (IL) tagset. Both tagged (600 sentences) and untagged (5000 sentences) are used for learning. The algorithm has achieved an accuracy of 92% for Gujarati texts where the training corpus is of 10,000 words and the test corpus is of 5,000 words.