Morphology analysis for Hidden Markov Model based Indonesian part-of-speech tagger
Muljono Muljono, Umriya Afini, Catur Supriyanto · 2017
Part-of-Speech (POS) tagging plays an important role in Natural Language Processing (NLP). It classifies a word into its tags, such as noun, verb, and pronoun. Many POS tagging approaches have been developed to solve manual POS tagging which is a time-consuming task. Hidden Markov Model (HMM) is a statistical-based method which widely used for POS tagging. In Indonesian language, HMM has been improved with affix tree method which handles Out-of-Vocabulary (OOV) words problem and affixation. The problem is affix tree does not provide any information to handle the clitics. Therefore, this study proposes morphology analysis for Indonesian Part-of-Speech (POS) Tagging. We combine MorphInd as morphology analyzer and HMM to improve the performance of POS tagging. In the experiment, there are 10,000 tokens for training and 3,000 tokens for testing. We prepare three different testing corpus; each consists of 10%, 20%, and 30% OOV words. The experimental results show that the proposed method achieves better performance compared to other methods.