Data Extraction and Integration from HTML Documents

Ou Hong Zhang Jian · Huadong Li-Gong Daxue xuebao · 2003

Using XML and HTML Tidy tools set, we can get a lightweight method of Web data mining and transformation. The purpose of transformation is to separate HTML document content from its schema. The processes included purifying HTML documents by HTML Tidy Standard class library, analyzing HTML element's structure through DOM, and extracting data with XSL and XPATH.

Read the paper · More papers on PaperTik