Finding the Number of Clusters in a Dataset Using an Information Theoretic Hierarchical Algorithm

Mohammad Aghagolzadeh, Hamid Soltanian‐Zadeh, Babak Nadjar Araabi, Ali Aghagolzadeh · 2006

One of the most challenging problems of clustering is detecting the exact number of clusters in a dataset. Most of the previous methods, presented to solve this problem, estimate the number of clusters with model based algorithms, which are not able to detect all types of clusters and also face a problem in detecting coupled clusters in a dataset. In this paper we propose a new method for finding the number of clusters in a dataset utilizing information theory and a top-down hierarchical clustering algorithm. The algorithm starts from a large number of clusters and reduces one cluster in any iteration and then allocates its data points to the remaining clusters. Finally, by measuring information potential, the exact number of clusters in a desired dataset is detected. Our method shows high capability and stability in detecting the number of clusters even in complex datasets, as it is computational efficient too. We show the effectiveness of the proposed method by experimenting on several artificial and real datasets and comparing its results with two developed methods for finding the number of clusters in a dataset. The comparisons show superiority of the proposed method.

Read the paper · More papers on PaperTik