Improving cross-document coreference
Octavian Popescu, Christian Girardi, Emanuele Pianta, Bernardo Magnini · 2008
In this paper we present a cross document coreference system which resolves some of the more problematic cases by taking into account pieces of evidence coming from different sources. The corpus we work with is a seven-year news collection from a local newspaper. The approach does not assume any prior knowledge about persons (e.g. an ontology) mentioned in the collection and requires basic linguistic processing (named entity recognition) and resources (a dictionary of person names). The system parameters have been estimated on a 5K corpus of Italian news documents. The evaluation, over a sample of four days news documents, shows that the error rate of the system (1.4%) is above a baseline (5.4%) for the task.