A statistical approach to errors in bibliographic data bases
Thomas Kuch · ACM SIGDOC Asterisk Journal of Computer Documentation · 1976
Virtually all large bibliographic data base (BDB) systems use inverted files for efficient retrieval. These files contain unique character strings found in the BDB, with the exception of stopwords such as A, AN, THE, etc. For each unique string, the number of records in which it occurs is also given.