Web document clustering algorithm based on semantic similarity

Jing Yang · Journal of Hefei University of Technology · 2009

A Web document clustering algorithm using semantic similarity(WDCSS) is presented in this paper.Firstly,the minimal spanning tree is obtained from the similarity of document .Secondly,the smallest similarity threshold is determined based on the analysis of probability and statistics.Finally,the minimal spanning tree is cut.Simultaneously,the subclasses are divided and merged.The experiment indicates that the WDCSS can not only accurately analyze the reasonable clusters and the exceptional samples from data sets with different cluster shapes,but it can also avoid the decreasing of the clustering quality influenced by the choice of users'parameters.

Read the paper · More papers on PaperTik