A Comparative Study on k-Means Clustering with Different Cluster Representations

Siyuan Li, Deqiang Han · 2023

K-means aims at partitioning a dataset into several clusters so that samples in the same cluster are compact and samples in different clusters are well separated. The centroid point with each dimension described by a real number, is used as a modeling approach for traditional k-means to represent the cluster, which is limited in information description. Fuzzy number contains more information compared with the real number for the point. Therefore, one can try to improve the clustering analysis performance using a richer representation of the cluster. In this paper, a comparative study among the different representations of each dimension of cluster including point, interval number, triangular fuzzy number and trapezoidal fuzzy number is conducted. Experimental results on both artificial datasets and real-world datasets demonstrate that using trapezoidal fuzzy number to represent cluster outperforms other representations with respect to four metrics: accuracy, normal mutual information, rand index, and adjusted rand index.

Read the paper · More papers on PaperTik