Research on a method of self-adaptation of the number of clusters for hierarchical initialization clustering
Wei Jian-don · Electronic Design Engineering · 2015
K-means algorithm is a common and effective clustering algorithm based on partition. To solve the problem of sensitivity of initial cluster centers, the most frequently used method is searching optimal initial cluster centers by hierarchically initializing. However, it also takes the number of clusters as the argument. It is so difficult to give the number of clusters for the high dimensional data and large volume data that the hierarchal initialization K-means cannot be directly applied. To address this problem, this paper proposes a Davies Bouldin Index(DBI) based hierarchical initialization K-means(DHIKM) algorithm through integrating DBI metric into hierarchical initialization K-means algorithm. By DBI metric, DHIKM can quickly determine the number clusters on sampled data. Experiments on UCI dataset and synthetic data demonstrate the effectiveness of the proposed algorithm.