Finding the Optimal Number of Clusters from Artificial Datasets

Niina Päivinen, Tapio Grönfors · 2006

This study deals with the problem of selecting the right number of clusters. Scale-free minimum spanning trees (SFMSTs) were constructed from the artificial test datasets, and the number of clusters, based on the distribution of the edge lengths, as well as the clustering itself was obtained from the structure. As a reference, the nearest neighbor and k-means clustering methods were used, and the number of clusters was determined with the largest average silhouette width criterium. The SFMST clustering mehtod proved to be a method which is able to automatically find the optimal number of clusters from the dataset without using any user-defined parameters.

Read the paper · More papers on PaperTik