The Extract of Web Site Structure & Web Page's Characteristic Information

Erping Wang · 2004

The World Wide Web is now emerging as a powerful global resource of rich information for web users. However,since the structure of hypertext is very complex and its cross-references are disoriented,it is very difficult to render them in a visual form. This paper briefly introduces the actuality of Web Visualization,and describes the technology of extracting structure and characteristic information from web pages,which is prerequisite of Web Visualization. It includes reading the contents of web pages,extracting hyperlinks embedded in web pages,transforming the hyperlinks in relative formats into the absolute ones,filtering out the redundant hyperlinks,defining and storing the tree structure etc.

Read the paper · More papers on PaperTik