Document image retrieval without OCRing using a video scanning system

Erçan E. Kuruoğlu, Vern T. Tan · 2000

In this paper, we propose a technique for efficient document retrieval from digital libraries containing document images which are token based compressed. The query image is captured from a paper document by the video scanning tool of a multimedia system. The technique we propose uses the layout information supplied by the relative positions of the character tokens on the page of a “query” paper document to retrieve the original document in the image database. This technique avoids OCRing the query document and the documents in the database; moreover avoids decompressing the token based compressed documents in the database, therefore achieving important time and computational gains.

Read the paper · More papers on PaperTik