Finding Structure and Characteristics of Web Documents for Classification.

Wai-ching Wong, Ada Wai-Chee Fu · 2000

Many Web documents containing the same type of information, would have similar structure. In this paper, we examine the problem of finding the structure of web documents and present a hierarchical structure to represent the relation among text data in the web documents. Due to the loose standard of web page publishing, different authors can use different wordings (labels) to label the same information. We introduced a labels discovery algorithm that uses the hierarchical structure extracted from the web pages. The algorithm discovers similar labels which describe the same kind of information. Such labels would help us find the structure of the web documents. Experiments have shown that the algorithm can successfully discover similar labels and the structure obtained by our method can distinguish web pages accurately. 1 Introduction The World Wide Web (WWW) provides a vast resource for information of almost all types. There are over half a billion homepages on the WWW and a thousand m...

Read the paper · More papers on PaperTik