An Improved Hierarchical K-Means Algorithm for Web Document Clustering
Yongxin Liu, Zhijng Liu · 2008
In order to conquer the major challenges of current Web document clustering, i.e. huge volume of documents, high dimensional process, we proposed a simple agglomerative hierarchical k-means clustering (SAHKC) algorithm based on H-K (hierarchical k-means) algorithm, and a new model was used in this paper to describe the Web document, named as multiple feature vector space model (MFVSM). Experimental results indicate that: the MFVSM is helpful in improving the quality of clustering result, and compare with the H-K algorithm, the SAHKC algorithmpsilas running time reduce nearly 30%, however, the average precision of clustering result only reduce about 10%.