Exploiting Lexical Dependencies from Large-Scale Data for Better Shift-Reduce Constituency Parsing

Muhua Zhu, Jingbo Zhu, Huizhen Wang · 2012

This paper proposes a method to improve shift-reduce constituency parsing by using lexical dependencies. The lexical dependency information is obtained from a large amount of auto-parsed data that is generated by a baseline shift-reduce parser on unlabeled data. We then incorporate a set of novel features defined on this information into the shift-reduce parsing model. The features can help to disambiguate action conflicts during decoding. Experimental results show that the new features achieve absolute improvements over a strong baseline by 0.9% and 1.1% on English and Chinese respectively. Moreover, the improved parser outperforms all previously reported shift-reduce constituency parsers. Title and Abstract in Chinese 利用大规模数据词汇依存关系改进移进-归约成分句法分析 本文提出了一种利用词汇依存关系改进移进-归约成分句法分析的方法。首先,我们利用 基准系统在大规模无标注数据上进行自动句法分析并从分析结果中抽取词汇依存关系。其 后,我们在词汇依存信息的基础上定义了一组新特征并将这些特征整合到移进-归约句法 分析模型 中。新特征用于帮助消除移进-归约过程中的动作歧义。实验结果表明,新特征 在英文和中文数据上分别取得了0.9% 和1.1%的性能改进。最终得到的句法分析器的性能 优于相关研究工作中所报告的移进-归约句法分析器的性能。

Read the paper · More papers on PaperTik