A Stochastic Part of Speech Tagger for the Sinhala Language based on Social Media Data mining

Shalki Ginthota Withanage, Thushari Silva · 2020

Part of Speech (POS) taggers plays a critical role in NLP applications for analyzing the behavior and the construction of a language. POS tagging is an essential step in NLP for further analysis of a language. Even though there are various POS taggers available for the English language, still there is no official POS tagger for the Sinhala language. To overcome this problem, this paper presents a stochastic based POS tagger for the Sinhala language based on social media data mining. This tagger uses a Hidden Markov model (HMM) with bi-gram probabilities for training and tagging data. The tagger is capable of handling lexical items with multiple POS tags while predicting the POS tags of unknown words. The Viterbi algorithm is used to decide the best POS tag for each word based on the results of HMM. An annotated corpus was developed using social media data during the research for HMM parameter estimation. The tagger has shown 63% overall accuracy and nearly 90% accuracy for known words. The proposed approach has demonstrated higher accuracy level compared to SVM and hybrid approaches.

Read the paper · More papers on PaperTik