Statistical models for case ambiguity resolution in Korean

Kihwang Lee · ERA · 2005

This thesis deals with the resolution of case ambiguity in Korean. Even though Korean is a case marked language, in which phonetically recognisable case markers (case particles) mark cases explicitly, nominal words without any accompanying case particles are used frequently in naturally occurring texts and speech. When the case particles are not present, it is basically a matter of conjecture to infer the grammatical function of the nominal words. The position of a nominal word itself cannot give much help as Korean is a relatively free word order language. The case ambiguity problem has brought a great controversy in Korean linguistics and has been regarded as an unavoidable obstacle for automatic processing of the Korean language. The aimof this thesis is to tackle the case ambiguity problem inKoreanwith statisticalmethods. To achieve the aim we pursue the following objectives. First, through an examination of the relevant theoretical work, we precisely define the realm of the case ambiguity problem in Korean. We also clearly identify the case particles that are involved in case ambiguitiy problem. Second, we clearly specify our knowledge-lean training data constructionmethod. We also attempt to measure the effectiveness of the data collectionmethod by applying the method to two treebanks of Korean. Third, we suggest two case decisionmethods for the task of case ambiguity resolution: discrete case decision method and sequential case decision method. In the discrete case decision method, each case ambiguity in a sentence is treated in isolation. For this method we use statistical classifiers based on simple joint probabilistic models that can be easily extended. We incorporate two new features, the list of neighbouring case particles and the distance between the focus nominal and the predicate, which have never been used in previous approaches. In the sequential case decision method, every case decision in a sentence is treated in the context of a series of case decisions that take place in the sentence. Thismethod is similar to other sequential category assignment tasks such as part-of-speech tagging. Thus we adopt the well-knownMarkov chain tagging model. Finally, our statistical case ambiguity resolutionmodels are evaluated by comparing the outputs of the system applied to a test set with the multiple human annotations on the test set. Kappa is used to measure the pairwise agreements between the system outputs and human annotations. From the evaluation results, we show the effectiveness of the two new features. As a conclusion, we present the contributions and the limitations of our approach to the case ambiguity problem. Several possible future directions are also laid out.

Read the paper · More papers on PaperTik