Improved Evaluation and Parse Reranking for Combinatory Categorial Grammar

Dominick Ng · 2010

Accurate syntactic parsing of text is crucial for automatically understanding language. This makes parser performance and the metrics used to measure parser performance central to the goals of natural language processing. Unfortunately, the field-standard parser evaluation methods are tied to particular grammar formalisms and testing data, and are uninformative when considering the cross-domain performance of a parser. The wide variety of parsers built on different formalisms makes it difficult to know whether new developments for improving parsing accuracy are relevant across different parsers. For example, supervised reranking of n-best parses has been shown to improve the accuracy of a number of parsers, but the potential benefits of reranking for parsers based on different grammars remains unclear. This thesis addresses several open questions regarding parser evaluation and performance. We assess a new style of parser evaluation for the C&C Combinatory Categorial Grammar parser, demonstrating that an extrinsic task-based evaluation of parsers is meaningful and can be successfully implemented. This improves on traditional evaluation metrics by being formalism-agnostic and requiring no expensive linguistic annotation to supply gold-standard test data. It also provides a fairer assessment of parser performance based on how parsing is used in the field. We develop a parser reranker for the C&C parser, and provide a comprehensive blueprint for implementing the features used by the system. Our initial experiments showed that state-of-the-art features from an existing reranker substantially reduce the performance of the C&C parser. Following a systematic analysis of errors made by the parser, we develop new reranking features that are better tailored to distinguishing cases of these errors. Our experiments showed that these features produce significant performance improvements over the baseline parser without the use of additional annotated data. Through this work we have shown that rerankers can produce respectable accuracy improvements for parsers if they are targeted towards the parser in question and the errors it makes. Insights from our analysis of reranking features and parser error will be useful for other reranker implementations, while our work on evaluation will contribute to more representative methods of evaluating parsing performance. The improved parser performance will lead to improvements in downstream systems, providing for more intelligent information search, storage, and management.

Read the paper · More papers on PaperTik