Automated information extraction from web pages using presentation and domain regularities

Srinivas Vadrevu · 2008

The vast amount of data on the World Wide Web poses many challenges in devising effective methodologies to search, access and integrate the information. Recently many information extraction systems have been proposed to (semi) automatically extract structured information from the Web. However, the applicability of these systems is limited only to a specific portion of the Web that is generated from underlying databases with a common presentation template. Some information extraction systems also require human assistance thereby limiting their scalability. Hence, the biggest challenge of automated information extraction from all kinds of Web documents still remains open. Even though the Web is composed of several heterogeneous sources, their data and metadata presentation adheres to certain regularities. This dissertation employs such presentation and metadata regularities to automatically extract structured information from Web data. A novel semantic partitioning information extraction algorithm is described that utilizes the presentation and domain metadata regularities to organize the content in a Web page into hierarchical group structures by using grammar induction techniques. This algorithm transforms Web pages into semantically represented XML (Extensible Markup Language) documents where the tag names denote the metadata labels and the elements denote the data value labels. The semi-structured XML documents that are extracted from Web pages facilitate efficient indexing of the content in Web pages by their metadata and can be used to provide structured query support over Web pages. It is shown how the semantic partitioning algorithm can be utilized to build a semantic search engine that can automatically reformulate the given keyword query into XQuery language that can be processed over the XML documents. The resulting search engine achieves high precision by retrieving the relevant fragments of the structured parts of the Web page as results instead of pointer URLs (Uniform Resource Locator) where the results can be found. This dissertation also discusses how the proposed semantic partitioning algorithm can be utilized to mine taxonomies and their instances from various collections of Web pages. The technical details of the algorithms are described and empirical evaluation performed on diverse collections of Web pages is presented to demonstrate their efficacy and scalability.

Read the paper · More papers on PaperTik