Part-of-Speech Tagging for Azerbaijani Language
Samir Mammadov, Samir Rustamov, Ali Mustafali, Ziyaddin Sadigov, Rasim Mollayev, Zamir Mammadov · 2018
The paper describes the process of implementing a HMM PoS tagger for Azerbaijani language to tag given text based on the tagged corpus. Different methodologies for part-of speech tagging have been studied, and after analysis of these methodologies, Hidden Markov Model has been chosen for implementation. For Azerbaijani language, the paper demonstrates the steps taken to build a stemmer as an essential part of PoS tagger. A thorough examination of possible word groups and exceptions has been conducted and most of such cases have been successfully handled. As of now, a tagged corpus of large enough size does not exist for Azerbaijani language and it hinders the testing process of HMM tagger. For this reason, a small corpus has been created for testing. However, as HMM shows remarkable performance when run on English corpus, it is expected that it will produce decent results for Azerbaijani language too.