Semantics of links and document structure discovery
John R. Punin, Mukkai S. Krishnamoorthy · 2002
This paper presents a novel algorithm to discover the hier-archical document structure by classifying the links between the document pages. This link classication adds metadata to the links that can be expressed using Resource Descrip-tion Framework Syntax [7]. Several well-known programs automatically generate HTML web pages from dierent doc-ument formats such as LaTeX, Powerpoint, Word, etc. Our interest is in the intertwined HTML web pages generated by the LaTeX2HTML program [6]. We use the web robot of the WWWPal System [11] to save the structure of the web document in a webgraph. Then the web analyzer of the sys-tem applies our algorithm to discover the semantics of the links and infer the hierarchical structure of the document. 1.