K-means Clustering Algorithm in Projected Spaces
Alissar Nasser, Denis Hamad, Chaïban Nasr · 2006
Clustering has been known as a popular technique for pattern recognition, image processing, and data mining. Unfortunately, all known clustering algorithms tend to break down in high dimensional spaces; this is due to the inherent sparsity of the points. We investigate, in this paper, the use of linear and nonlinear principal manifolds for learning low-dimensional representations for clustering. Several leading methods: PCA, KPCA, Sammon, and CCA are examined and tested in clustering experiments using synthetic and real datasets from the UCI databases. We compare the clustering performance of the K-means algorithm on data projected by these projection methods. The experimental results show that K-means clustering on data projected by KPCA outperforms those projected by the three other methods