Study on Web Data Extraction Based on XML
Yu Li · Computer Technology and Development · 2007
Introduces three common methods for Web data extraction:method that directly analyses HTML document,method that bases on XML(it is also called method that analyses the structure of HTML document)and conceptual-model-based approach.especially,Web data extraction based on XML is studied.The original HTML document gets through a filter which checks and corrects the syntax structure of HTML document,then forms an well-formed XHTML,XML stools can be used to dispose these HTML documents.Implemented a data preprocessing which transformed the semi-structured HTML document to the structured XML document.Also it created a good condition of using the traditional data extraction methods to deeply data extraction.