Statistical Tagger for Bhojpuri (employing Support Vector Machine)
Srishti Singh, Girish Nath Jha · 2015
The authors present the first Support Vector Machines (SVM) based statistical Parts of Speech (POS) Tagger developed for Bhojpuri. Bhojpuri is a less resourced Indo Aryan language of the Asian continent and the POS tagger presented here is a step towards developing language resources for it. SVMs have already been trained on other languages like Malayalam and Bengali with an accuracy of 86-90 %. The present research came up with approximately 87.3 -88.6% accuracy for test datasets.