Development of Pairwise Comparison-based Japanese Dependency Parsers and Application to Corpus Annotation

Masakazu Iwatate · Institutional Repositories DataBase (IRDB) · 2012

For machine-learning-based Japanese dependency parsing, it is essential to develop an accurate parsing model.It is also essential to construct a corpus of a new domain efficiently when parsing sentences of the domain.For developing an accurate parsing model, we propose three approaches: developing a single accurate model, combining parsers, and combining a parser with a coordination analyzer.Considering selectional preferences between a dependent bunsetsu and its candidate head bunsetsus plays an important role to construct an accurate parser.In Japanese dependency parsing, Kudo's relative preference-based model [27] outperforms both deterministic parsing models and probabilistic CFG-based parsing models.In the relative preference-based model, an ME model estimates selectional preferences for all candidate heads, which cannot be considered in the previous parsing models.We propose a parsing model in which the selectional preferences are directly modeled by one-on-one games in a step-ladder tournament.In the evaluation experiment with Kyoto Text Corpus Version 4.0, the proposed model outperforms the previous researches, including the relative preference-based model.We also investigates the accuracy of partial parsing of three SVM-based Japanese dependency parsing models: the shiftreduce model, the cascaded chunking model, and the tournament model.We show the coverage-accuracy curves based on the scores produced by SVMs.Our performance evaluation with the Kyoto Text Corpus shows that the partial parsing accuracy of the tournament model is the highest among the three dependency parsing models.The tournament model achieves 99% accuracy with 60% bunsetsu-based coverage.

Read the paper · More papers on PaperTik