Improving Suffix Tree Clustering Algorithm for Web Documents
Zhuang Yan, Youguang Chen · Advances in intelligent systems research/Advances in Intelligent Systems Research · 2015
Web document clustering results can help users quickly locate the information they need among the results search engines returned.According to the characteristics of the suffix tree structure and the flaws of similarity calculation in STC algorithm's cluster merging, this paper proposes an improved suffix tree clustering method.The method combines vector space model with Pearson correlation coefficient, calculates the relevant of clusters based on document vector of all clusters, and then utilizes the relevant vectors of clusters and the correlations between them to calculate the similarity for cluster merging, improves the clustering process of documents.Analysis of the experimental results shows that the method outperforms the original STC algorithm on Web documents clustering.