Computerized scheme for duplicate checking of bibliographic data bases

C.A. Giles, Alex Brooks, Tamas E. Doszkocs, D.J. Hummel · OSTI OAI (U.S. Department of Energy Office of Scientific and Technical Information) · 1976

A technique for the automatic identification of duplicate documents within large bibliographic data bases has been designed and tested with encouraging results. The procedure is based on the generation and comparison of significant elements compressed from existing document descriptions. Problems arising from inconsistencies in editorial style and data base formats and from discrepancies in spelling, punctuation, translation and transliteration schemes are discussed; one method for circumventing ambiguities and errors of this type is proposed. The generalized computer program employs a key-making, sorting, weighting, and summation scheme for the detection of duplicates and, according to preliminary findings, achieves this objective with a high degree of accuracy. Sample results from five large data bases suggest that this automatic system performs as effectively as manual techniques.

Read the paper · More papers on PaperTik