Three-Step Probabilistic Model for Korean Morphological Analysis
Jae Sung Lee · Jeongbo gwahaghoe nonmunji. so'peuteuweeo mich eung'yong · 2011
A morphological analyzer based on probabilistic model can learn easily various language phenomena and tagging principles used in morpheme-tagged corpus, so that it is very portable to various domains. In this paper, we propose a three-step probabilistic model for Korean morphological analysis which consists of original form restoring step, morpheme segmentation step and morpheme tagging step. The three-step method, which uses modular approach, reduces processing complexity compared with two-step probabilistic model. Processing in Jaso unit rather than syllable unit and using morpheme transition probahility for morpheme segmentation increase portahility for various tagging principles. Experiment on Sejong tagged corpus, both of written text corpus and spoken text corpus, was done to show the performance of the model and compare it with other methods.