Intelligent Wrapping from PDF Documents

Tamir Hassan, Robert Baumgartner · 2005

Wrapping is the process of navigating a data source, semi- automatically extracting data and transforming it into a form suitable for data processing applications. The semi-structured form of web pages, coupled with the availability of business-relevant data, has led to the availability of several established products on the market for wrapping data from the Web. One such approach is the Lixto methodology (1), a result of research performed at DBAI. Many commercial applications also require the extraction of data from PDF documents. There appear to be no general-purpose approaches to fulfil this need and, as the PDF format is unstructured, this is a challeng- ing task. We are investigating PDF data extraction in the NEXTWRAP project. This paper presents our work in progress, with particular refer- ence to low-level segmentation algorithms.

Read the paper · More papers on PaperTik