Hybrid Part of Speech tagger for Sinhala Language

D. A. N. Gunasekara, W.V. Welgama, A. R. Weerasinghe · 2016

This research presents a hybrid Part of Speech tagging approach which utilizes both rule based and stochastic tagging approaches for Sinhala Language. In the first phase, Hidden Markov Model based stochastic tagger is constructed which is based on bi-gram probabilities. A stemmer is used in the tagging process to enhance the accuracy of the tagger. An experiment on three POS tag set versions is carried out to come up with the best tag set which leads towards a meaningful and precise tagging process for Sinhala Language. Since Sinhala is a morphologically rich language, rules based on morphological features are used to predict the relevant tag for words which do not present in the training set. Further, an experiment is carried out to find out whether the implemented hybrid tagger can be used to enhance the size of the data set. The implemented hybrid tagger is successful in achieving an overall accuracy of 72% when the average unknown word percentage is 20%.

Read the paper · More papers on PaperTik