The Effectiveness of Using Malay Affixes for Handling Unknown Words In Unsupervised HMM POS Tagger

Hassan Mohamed, Nazlia Omar, Mohd Juzaiddin Ab Aziz · International Journal of Engineering & Technology · 2018

The challenge in unsupervised Hidden Markov Model (HMM) training for a POS tagger is that the training depends on an untagged corpus; the only supervised data limiting possible tagging of words is a dictionary. A morpheme-based POS guessing algorithm has been introduced to assign unknown words’ probable tags based on linguistically meaningful affixes. Therefore, the exact morphemes of prefixes, suffixes and circumfixes in the agglutinative Malay language is examined before giving tags to unknown words. The algorithm has been integrated into HMM tagger which uses HMM trained parameters for tagging new sentences. However, for unknown words their parameters are absent. Therefore, the algorithm applies two methods for assigning unknown words’ emission to HMM tagger, first is based on uniform distribution of all possible tags; and second, is based on marginal proportionate distribution of tags. The effective method is proven to be using morpheme-based POS guessing with unknown word emissions substituted by a value proportionate to the marginal distribution of tags.

Read the paper · More papers on PaperTik