Training Dependency Parsers from Partially Annotated Corpora

Daniel J. Flannery, Yusuke Miayo, Graham Neubig, Shinsuke Mori · International Joint Conference on Natural Language Processing · 2011

We introduce a maximum spanning tree (MST) dependency parser that can be trained from partially annotated corpora, allowing for effective use of available linguistic resources and reduction of the costs of preparing new training data. This is especially important for domain adaptation in a real-world situation. We use a pointwise approach where each edge in the dependency tree for a sentence is estimated independently. Experiments on Japanese dependency parsing show that this approach allows for rapid training and achieves accuracy comparable to state-ofthe-art dependency parsers trained on fully annotated data.

Read the paper · More papers on PaperTik