An Improved K_Means Algorithm for Document Clustering Based on Knowledge Graphs
Xiaoli Wang, Ying Li, Meihong Wang, Zixiang Yang, Huailin Dong · 2018
K _means algorithm is one of the typical clustering algorithms in text mining tasks. K_means algorithm is widely used in many areas because of its easy to implement and ability to handle large datasets with better scalability. However, the random selection of initial cluster centroid in traditional K_means algorithm for text clustering easily leads to local optimization and instability of clustering results. Therefore, in order to overcome this shortcoming, this paper propose an improved K_means algorithm for document clustering which based on following two points: (i)we used concept distance to optimize the choice of the initial cluster centroid, which can avoid the drawbacks caused by random selection; (ii)we adopted knowledge graphs to improve traditional k_means text clustering algorithm by optimizing the calculation of text similarity. Theoretical analysis and experimental results show that the improved algorithm could optimize the accuracy of text clustering effectively.