Prediction system for Sindhi parts of speech tags by using support vector machine
Farhan Ali Surahio, Javed Ahmed Mahar · 2018 International Conference on Computing, Mathematics and Engineering Technologies (iCoMET) · 2018
Language is the resource of communication and each language has some grammatical rules. Sindhi is also a language and being considered as an ancient language that follows some Parts of Speech (POS) rules escort with tags or entire words. It is the method of giving a task to proper speech type or to a quantity of expression in language processing pronunciations. A programmed explanation of POS tag presented in this paper based on the Sindhi Text. To do this, 28000 words corpus has collected covering, Poetry text taken from primary text books, newspapers and stories etc. Similarly, amount of words have separated into various size files for testing and training purposes. In our approach processing mechanism Support Vector Machine (SVM) is used to tag the sentences of Sindhi language. According to our perception, this approach has never been used to tag Sindhi sentence and result would display the accuracy of SVM tagger better than the exits one approaches. In the existing POS tagger system, an accuracy of POS tagging for unspecified words is less than for recognized words and sentences. But, in our proposed tagging system found better accuracy for ambiguous and unidentified tagged words. An accuracy of 97.86% received with our proposed system tagger which is better than the presented approaches we believed.