WEB STRUCTURE ANALYSIS FOR INFORMATION MINING
Vijjappu Lakshmi, Ah‐Hwee Tan, Chew-Lim Tan · Series in machine perception and artificial intelligence · 2003
Our approach to extracting information from the web analyzes the structural content of web pages through exploiting the latent information given by HTML tags. For each specific extraction task, an object model is created consisting of the salient fields to be extracted and the corresponding extraction rules based on a library of HTML parsing functions. We derive extraction rules for both single-slot and multiple-slot extraction tasks which we illustrate through two sample domains.