Text Snippets from the DomGraph
A. Ratkiewicz, Filippo Menczer · 2008
We explore the idea that the Document-Object Model tree of an HTML page | absent any semantic or heuristic interpretations of the tags and their positions | provides cues about the importance of the information it contains. This hypothesis is evaluated by constructing a DomGraph, i.e., a network of the DOM trees of pages connected by their hyperlinks, and using it in a snippet-extraction technique. In this process, we also address technical issues related to the processing of the resultant very large graph. Snippets produced by this technique are compared in a user study to those extracted by a reasonable and simple baseline method, and found to be clearly preferred by users.