A New Initialization Method for K-means Clustering
Elahe Baratalipour, Seyed Jahanshah Kabudian, Z. Fathi · 2024
Clustering techniques are widely used for data analysis and knowledge discovery in various fields. Among them, the K-means algorithm is a popular and efficient method for categorization of data into clusters based on similarity. However, the effectiveness of the K-means algorithm heavily depends on the initial selection of cluster centroids. In traditional K-means, the initial centroids are randomly chosen from the input data. This random initialization can lead to suboptimal clustering results, as the algorithm is sensitive to the initial configuration. The resulting clusters may not accurately represent the underlying structure of the data and may vary for different runs of the algorithm. In the proposed improved K-means for the initial centroids, we first find the largest and smallest value based on top-two important features of the data set and perform clustering in this space. Then we take the data close to the average of each cluster as the center of that cluster. The experimental results on three well-known datasets, TIMIT, Iris and Immunotherapy shows that the proposed initialization method outperforms other ones in terms of clustering validity and accuracy measures.