The XML-based Information Extraction on Data-intensive Page

Yanheng Li · 2007 IFIP International Conference on Network and Parallel Computing Workshops (NPC 2007) · 2007

This paper puts forward an XML-based information extraction method which applies XSLT and XPath technology to construct extraction rules. The aim of this method is to extract useful information from data-intensive pages. This paper firstly analyzes the traits of data- intensive pages. Aiming at those traits, we proposed a path induction method to conclude record pattern of pages, to obtain the path expression of useful information, and eventually to construct extraction rules. Furthermore, this paper presents the method of optimization of extraction rules in order to getting more robust rules.

Read the paper · More papers on PaperTik