Machine Learning-based approach to automatic POS tagging of Macedonian language

Martin Bonchanoski, Katerina Zdravkova · 2017

This paper presents the research that has contributed to the creation of an automatic part-of-speech (POS) tagger of Macedonian, a Slavic language that has a rich morphology, but limited language resources and contributions towards establishing of Natural Language Processing (NLP) tools. The created system automatically tags only the part-of-speech category for each word, without giving the full annotation. For the purposes of this research, the existing large online lexicon was combined together with a self-created crowd-sourcing system intended for manual disambiguation of POS tags. The joint technique resulted in a POS tagged corpus that was used during the training and testing phase. Four different models were produced, built using: TnT tagger, averaged perceptron, cyclic dependency network and guided learning framework for bidirectional sequence classification and the performance of all the models was exhaustively compared. The final accuracy that has been achieved is 96.37%, reaching a result which is comparable to more researched languages.

Read the paper · More papers on PaperTik