Hierarchical clustering of Chinese web pages based on suffix tree
Chao Ke · Journal of Liaoning Technical University · 2006
In order to facilitate users browsing web search results produced by search engines,a new method called STCC algorithm is proposed,which combines STC algorithm and chameleon algorithm to group similar Chinese web pages in a hierarchical fashion.This method employs Jaccard coefficient to modify the similarity measure of base cluster in STC,then according to the similarity matrix of base cluster,chameleon algorithm is used to cluster web pages.Experimental results show that the precision in STCC increases by nearly ten percent compared with that in STC,meanwhile,chain effect in single-link algorithm can be avoided by using STCC algorithm,and it is suitable for large scale web pages clustering.