A Comparison Study of Cluster Validity Indices Using a Nonhierarchical Clustering Algorithm

Yosung Shim, Jiwon Chung, Inchan Choi · 2006

Cluster analysis is widely used in the initial stages of data analysis and data reduction. The K-means algorithm, a nonhierarchical clustering algorithm, has regained popularity among researchers in data mining and knowledge discovery, partly because of its low time complexity. The algorithm requires the number of clusters as an input parameter. When the parameter value is not known a priori, a researcher often has to use a cluster validity index to search for a suitable parameter value. In this study, we use computational experiments to examine the performance of cluster validity indices with the K-means algorithm. Our analysis parallels the study performed by Milligan and Cooper on cluster validity indices; we use hierarchical clustering algorithms and present observations and conclusions resulting from the simulation study... ..

Read the paper · More papers on PaperTik