Integrating Dictionaries into an Unsupervised Model for Myanmar Word Segmentation

Ye Kyaw Thu, Andrew Finch, Eichiro SUMITA, Yoshinori Sagisaka · 2014

This paper addresses the problem of word segmentation for low resource languages, with the main focus being on Myanmar language.In our proposed method, we focus on exploiting limited amounts of dictionary resource, in an attempt to improve the segmentation quality of an unsupervised word segmenter.Three models are proposed.In the first, a set of dictionaries (separate dictionaries for different classes of words) are directly introduced into the generative model.In the second, a language model was built from the dictionaries, and the n-gram model was inserted into the generative model.This model was expected to model words that did not occur in the training data.The third model was a combination of the previous two models.We evaluated our approach on a corpus of manually annotated data.Our results show that the proposed methods are able to improve over a fully unsupervised baseline system.The best of our systems improved the F-score from 0.48 to 0.66.In addition to segmenting the data, one proposed method is also able to partially label the segmented corpus with POS tags.We found that these labels were approximately 66% accurate.

Read the paper · More papers on PaperTik