Framework of web page analysis and content extraction with coordinate trees

Huamin Feng · Journal of Tsinghua University(Science and Technology) · 2005

With the development of the internet, the information on the Web has exploded and the presentations of the Web documents are also diverse, the latter brings tremendous troubles to the computer processing, such as information retrieval, information extraction etc. Considering the semi-structure of HTML document and lack of position information of the DOM tree, this paper presents a new framework of web page analysis and content extraction, which includes a novel coordinate tree model containing position information and a graph model reflecting the spatial relations. By transforming HTML documents into coordinate trees, the Web pages were analyzed and contents were extracted based upon the features of position and spatial relations. Experiment results on a set of 5 000 Web pages from 120 different sites show that the approach can achieve 93.78% in accuracy.

Read the paper · More papers on PaperTik