Noun phrase coreference resolution: a knowledge-rich cluster-based approach

Vincent Ng, Altaf Rahman · 2012

Noun phrase (NP) coreference resolution is the task of determining which NPs in a text or dialogue refer to the same real-world entity. Despite the fact that coreference resolution has been an active area of research in the NLP community for more than 40 years, it is far from being solved. The state-of-the-art in coreference resolution is rather unsatisfactory for at least three reasons. First, there is a lack of an accurate computational model for coreference resolution. Currently, the most commonly adopted approach for NP coreference resolution is the supervised machine learning approach, where a coreference classifier is trained from coreference-annotated data to determine whether two NPs in a text are coreferent or not. This mention-pair model suffers from two major weaknesses. First, since each candidate antecedent for an NP to be resolved is considered independently of the others. Second, the model is not expressive: the information extracted from the two NPs alone is not sufficient to mak an informed decision. Second, the majority of the existing coreference systems do not have access to the sophisticated knowledge sources needed for accurate resolution. Many coreference relations can only be identified using world knowledge. Third, the extensive use of supervised learning techniques for English coreference resolution presents a challenge to deploying coreference technologies to other languages. Specifically, for each new language of interest, one has to go through the labor-intensive, time-consuming process of hand-annotating documents. The goal of this dissertation is to advance the state-of-the-art in coreference resolution by aforementioned issues. In particular, we make the following contributions. First, we propose a new supervised coreference model, the cluster-ranking model, which addresses the weaknesses of the mention-pair model. In addition, we propose several linguistic and extra-linguistic extensions to the model to further improve its performance. Second, we examine the resolution of complex cases of definite pronouns. We propose new linguistic features that encode world knowledge specifically for resolving these difficult pronouns. Finally, we propose a translation-based projection approach to multilingual coreference resolution to combine online machine translation services with projection techniques to automatically project coreference annotations from a resource-rich language to the new language.

Read the paper · More papers on PaperTik