Constructing a Practical Constituent Parser from a Japanese Treebank with Function Labels

Takaaki Tanaka, Masaaki Nagata · 2013

We present an empirical study on constructing a Japanese constituent parser, which can output function labels to deal with more detailed syntactic information.Japanese syntactic parse trees are usually represented as unlabeled dependency structure between bunsetsu chunks, however, such expression is insufficient to uncover the syntactic information about distinction between complements and adjuncts and coordination structure, which is required for practical applications such as syntactic reordering of machine translation.We describe a preliminary effort on constructing a Japanese constituent parser by a Penn Treebank style treebank semi-automatically made from a dependency-based corpus.The evaluations show the parser trained on the treebank has comparable bracketing accuracy as conventional bunsetsu-based parsers, and can output such function labels as the grammatical role of the argument and the type of adnominal phrases.

Read the paper · More papers on PaperTik