Statistical Korean dependency parsing model based on the surface contextual information
Hoo-Jung Chung · 2004
Natural language parsing is a key problem to many tasks that require natural language processing. Many language processing tasks use the information on predicate-argument relation or modifier-modifyee relation and the parsing makes the extraction of the informa-tion possible by identifying relations between words, or phrases in sentences. However, it is difficult to parse a sentence correctly, because of the ambiguity inherent in the natural language. During the last decade, the statistical approach becomes the major trends in natural language parsing, or syntactic disambiguation. The two most important things in the statistical natural language parsing is selecting appropriate features that help syntactic disambiguation and designing a statistical model using them. This dissertation argues that the influence of surface contextual information, such as modification distance and local context, in solving syntactic ambiguity of the Korean lan-guage, and proposes the parsing model that considers a modification distance in a certain local context in addition to the preference for lexical bigram dependency. All of these pref-erences are expressed by probabilities conditioned on local context. The parsing model is