Methods for precise named entity matching in digital collections
Peter T. Davis, David K. Elson, Judith L. Klavans · 2003
In this paper, we describe an interactive system, built within the context of CLiMB project, which permits a user to locate the occurrences of named entities within a given text. The named entity tool was developed to identify references to a single art object (e.g. a particular building) with high precision in text related to images of that object in a digital collection. We start with an authoritative list of art objects, and seek to match variants of these named entities in related text. Our approach is to “decay ” entities into progressively more general variants while retaining high precision. As variants become more general, and thus more ambiguous, we propose methods to disambiguate intermediate results. Our results will be used to select records into which automatically generated metadata will be loaded. 1. Computational Linguistics and Metadata CLiMB (Computational Linguistics for Metadata Building, 1 funded by the Mellon Foundation) is an interdisciplinary project that aims to improve access to scholarly digital image collections by extracting descriptive metadata about images from related texts. With the large volume of image collections now being scanned, it is prohibitively expensive for specialized image catalogers to manually assign robust metadata to every image. By analyzing scholarly texts and associating their contents to related images, CLiMB tools will explore the potential to identify descriptive metadata which can be used to enrich catalog records. To test the process, these records will be mounted in a standard retrieval platform, where users will search for images related to particular keywords generated by CLiMB. Although the current CLiMB project is aimed at text associated with the information in image collections, our techniques and tools are applicable to texts of many types, not 1