Structured Information Extraction System from Web Pages

B. Anantha Barathi · 2014

The World Wide Web is a vast and rapidly growing source of information. Most of this information is in the form of unstructured text, making information hard to query. Many websites that have large collections of pages containing structured data. For example, Amazon lay out the author, title, comments, etc. in the same way in all its book pages. The values used to generate the pages (e.g. the author, title,..) typically come from a database. In this paper, we study the problem of automatically extracting the database values from template generated web pages without learning examples or other similar human input. We formally define a template and propose a model that describes how values are encoded into pages using a template. We present an algorithm that takes, as input, a set of template-generated pages, deduces the unknown template used to generate the pages and extracts, as output, the values encoded in the pages. Experimental evaluation on a large number of real input page collections indicates that our algorithm correctly extracts data in most cases.

Read the paper · More papers on PaperTik