Quantitative evaluation of coreference algorithms in an information extraction system
Robert Gaizauskas, Kevin Humphreys · Studies in corpus linguistics · 2000
Algorithms for performing coreference resolution can only be precisely evaluated given a benchmark corpus of coreference-annotated texts, together with techniques for evaluating the algorithms' output against the corpus. Such a corpus and such techniques have become available for the first time as part of the Message Understanding Conference 6 (MUC-6) evaluations of information extraction systems. In this paper we describe the MUC-6 coreference task and the approach to taken to it by the Large Scale Information Extraction (LaSIE) system developed at the University of Sheffield. The basic coreference algorithm used by this system is described in detail, as well as a set of variants, which allow us to experiment with different constraints such as restrictions to certain classes of anaphor, distance restrictions between anaphor and antecedent, and weighting factors in assessing semantic similarity of potential coreferents. Quantitative evaluation results are presented for these variants, ...