Automatic Data Extraction from Template Generated Web Pages.
Ling Ma, Nazli Goharian, Abdur Chowdhury · 2003
Information Retrieval calls for accurate web page data extraction. To enhance retrieval precision, irrelevant data such as navigational bar and advertisement should be identified and removed prior to indexing. We propose a novel approach that identifies the web page templates and extracts the unstructured data. Our experimental results on several different web sites demonstrate the feasibility of our approach.