Improving cross-document coreference

Octavian Popescu, Christian Girardi, Emanuele Pianta, Bernardo Magnini · 2008

In this paper we present a cross document coreference system which resolves some of the more problematic cases by taking into account pieces of evidence coming from different sources. The corpus we work with is a seven-year news collection from a local newspaper. The approach does not assume any prior knowledge about persons (e.g. an ontology) mentioned in the collection and requires basic linguistic processing (named entity recognition) and resources (a dictionary of person names). The system parameters have been estimated on a 5K corpus of Italian news documents. The evaluation, over a sample of four days news documents, shows that the error rate of the system (1.4%) is above a baseline (5.4%) for the task.

Read the paper · More papers on PaperTik