Data Extraction and Integration from HTML Documents
Ou Hong Zhang Jian · Huadong Li-Gong Daxue xuebao · 2003
Using XML and HTML Tidy tools set, we can get a lightweight method of Web data mining and transformation. The purpose of transformation is to separate HTML document content from its schema. The processes included purifying HTML documents by HTML Tidy Standard class library, analyzing HTML element's structure through DOM, and extracting data with XSL and XPATH.