A Clustering Algorithm for Automatically Determining the Number of Clusters Based on Coefficient of Variation
Tengteng Liu, Shouning Qu, Kun Zhang · 2018
The k-means algorithm is a typical clustering algorithm based on partition. The k-means++ algorithm is a high-quality clustering algorithm, and it is used to solve the problem that the traditional k-means algorithm is sensitive to initial centers. However, the original k-means++ algorithm is sensitive to outliers and needs to manually set the number of clusters. We propose an improved k-means++ clustering algorithm that automatically determine the number of clusters based on coefficient of variation, named CV-means++. Firstly, we propose a method to confirm initial centers by using density index of data points to avoid selection of abnormal data. Secondly, we introduce the concept of coefficient of variation, and calculate the relationship between the average intra-cluster coefficient of variation and the smallest inter-cluster coefficient of variation of k+(k+ > k) clusters to determine whether the number of clusters is optimal. Experiments performed on the UCI datasets demonstrate effectiveness of the algorithm.