A complex document information processing prototype

Shlomo Argamon, Gady Agam, Ophir Frieder, D. Grossman, David M Lewis, Gunho Sohn, KJ Voorhees · 2006

We developed a prototype for integrated retrieval and aggregation of diverse information contained in scanned paper documents. Such complex document information processing combines several forms of image processing together with textual/linguistic processing to enable effective analysis of complex document collections, a necessity for a wide range of applications. This is the first system to attempt integrated retrieval from complex documents; we report its current capabilities.

Read the paper · More papers on PaperTik