K-means Clustering for Handling High-dimensional Large Data
Tae‐Sik Yoon, Kyuseok Shim · 2012
Clustering is the one of the most popular data mining algorithms, which groups similar points in the same cluster while putting dissimilar points in different clusters. In this paper we propose efficient k-means clustering algorithms for high-dimensional large data. The proposed method utilizes the precomputed distances between points and dormant clusters to reduce distance computations. As the result of experiments, our methods outperform the existing work with respect to the total running time. Moreover, the proposed method uses less space compared to the previous works.