Modeling Unrestricted Coreference in OntoNotes
Sameer Pradhan, Lance Ramshaw, Mitchell Marcus, Martha Stone Palmer, Ralph Weischedel, Nianwen Xue · 2011
The importance of coreference resolution for the entity/event detection task, namely identifying all mentions of entities and events in text and clustering them into equivalence classes, has been well recognized in the natural language processing community. Automatic identification of coreferring entities and events in text has been a uphill battle for several decades, partly because it can require world knowledge which is not well-defined. Current techniques rely primarily on surface level features such as string match, proximity, and edit distance and on shallow linguistic features such as number, gender, Hobbs distance, etc. In the past researchers have tried using ontologies such as WordNet to extract useful features, and in the more recent past, there have been successful attempts to utilize information from much larger albeit noisier resources such as Wikipedia. There seems to be a growing consensus among researchers that some form of joint inference is a necessary next step towards improving the current state of the art. The computational learning community is also witnessing a move towards evaluations based on joint inference, with the past two tasks devoted to joint learning of semantic dependencies. One principle ingredient for joint learning is the presence of multiple layers of semantic information. The de facto standard datasets for current coreference studies, the MUC and the ACE corpora, on the other hand, are limited both in the size and in layers of information represented,