Korean Dependency Parsing Based on Machine Learning of Feature Weights

임수종, Young Tae Kim, Dong-Yul Ra · Jeongbo gwahaghoe nonmunji. so'peuteuweeo mich eung'yong · 2011

In this paper, we introduce a method for Korean dependency parsing based on machine learning by learning and using feature weights. The set of features of the system is constructed by generating a given number of features for every possible dependency relation. The degree of importance of a feature is represented by its weight. The weights are learned by using a training corpus in which dependency relations are tagged for each sentence. For this purpose, we used Sejong corpus tagged with phrase-structure trees and the ETRI corpus tagged with dependency structures. The training method we adopted is an on-line learning algorithm. It exploits a technique similar to that of the MIRA learning algorithm which is based upon the concept of max-margin. The experimental results showed that our parsing system's performance was 88.15% of dependency relation accuracy when developed with Sejong corpus and 88.06% of accuracy when developed with ETRI corpus. This demonstrates that our method allows a development of a Korean dependency parsing system with high performance.

Read the paper · More papers on PaperTik