Rule-based Approach to Korean Morphological Disambiguation Supported by Statistical Method

Minjung Kim, Hyuk‐Chul Kwon, Ae-Sun Yoon · Institutional Repositories DataBase (IRDB) · 1996

Korean as an agglutinative language shows its proper types of difficulties in morphological disambiguation, since a large number of its ambiguities comes from the stemming while most of ambiguities in French or English are related to the categorization of a morpheme.The current Korean morphological disambiguation systems adopt mainly statistical methods and some of them use rules in the postprocess.In our approach, the morphological analyzer reduces the number of the candidate morpheme strings using adjacency conditions when it analyses a word into morpheme strings.And then the disambiguation depends on rules and statistics successively.As for the rules, the partial parsing using finite state automata decides the compatibility of each pair of words: a negative value is assigned if a word can not co-occur with another word, while a positive value is given if they are compatible.After applying all the rules related to the word, our system chooses only the positively valued strings.When more than two strings still have same value, the priority in the context is decided by the statistics in the next stage.The accuracy of our approach as Korean tagging system is about 97.1% and it may yeild a better result than the Korean morphological disambiguation systems.

Read the paper · More papers on PaperTik