Improved K-means Based on Density Parameters and Normalized Distance
Xing Che, HengYi Tao, ZiHan Shi · 2021
The existing K-means has problems such as random selection of initial cluster centers, sensitivity to outliers, and inability to unify the data size. In this paper we describe a new improved K-means based on density parameters and normalized distance (K-DPND) we developed that addresses these problems. In the stage of selecting the initial cluster centers, K-DPND constructs a density parameter set according to the distance matrix and the average distance of the data set, selects the largest density parameter point as the cluster center, sets the density parameter of the point whose distance from this cluster center is less than the average distance to 0. Then loops until k initial cluster centers are found. In the clustering stage, the normalized distance is used to replace the Euclidean distance to calculate to determine the cluster to which each point belongs, and the median is used to replace the mean to calculate the new cluster center. Finally, compared with K-means, K-mediods, IK-DM and KICIC, K-DPND has good clustering results in most cases.