Report on the TREC-5 Confusion Track

Paul B. Kantor, Ellen M. Voorhees · 1996

For TREC-5, retrieval from corrupted data was studied through retrieval of single target documents from a corpus which was corrupted by producing page images, corrupting the bit maps, and applying OCR techniques to the results. In general, methods which attempted a probabilistic estimation of the original clean text fare better than methods which simply accept corrupted versions of the query text. 1 History The confusion track originated at an informal meeting held during TREC-3, stimulated by interest in the potential of various schemes for retrieval based on imperfect OCR applied to scanned legacy texts. The guiding idea was that even an imperfect translation of the image into text might support effective retrieval, especially if the retrieval were based on text representations not dependent upon the identification of terms. Participants in this first meeting included Mark Damashek (NSA), David Grossman (GMU), Fritz Nordby (then at Paracel), and Paul Kantor, who was selected, on the...

Read the paper · More papers on PaperTik