K-means text clustering algorithm based on initial cluster centers selection according to maximum distance
Zhai Dong-ha · Jisuanji yingyong yanjiu · 2014
Due to the random selection of initial cluster centers,K-means clustering algorithm is prone to local optimal and instability of clustering results,and huge number of iterations. To overcome the above problems,this paper selected the initial cluster centers according to maximum distance,and it was based on the fact that the farthest samples were the least likely in the same cluster. To apply the improved algorithm into text clustering,it constructed a method to transform text similarity into text distance,and also reconstructed cluster center iteration formula and measurement function. It employed a text set which included 5 categories and 1 500 texts in the experiment. The experimental results show that,compared with the original Kmeans algorithm and its two recently improved editions,the proposed method can improve the F-measure and reduce total consuming time.