Research on the Method of Clustering Web Documents Based on Hyperlink Information

Lina Sun · Computer Knowledge and Technology · 2006

Facing the massive volume text data information, how to locate the required information is one of the important research directions of text mining. The algorithms of text classification and clustering are applied to information retrieval, so the method of clustering Web documents based on hyperlink is presented according to the especial feature. And then the topological structure of website are found through hyperlink information, those noise and surplus hyperlink are cut down, the clusters are carried out based on the similarity between characteristic vectors which get from the content excavate of hyperlink anchor texts and web page texts. At the same time, the cluster centurions are adjusted dynamically, so as to realize the Web documents clustering based on hyperlink.

Read the paper · More papers on PaperTik