Initial Seeds Selection for K-means Clustering Based on Outlier Detection

Zhiyong Yang, Feng Xian Jiang, Xu Yu, Junwei Du · 2022

K-means clustering is a widely used algorithm in cluster analysis. However, the selection of initial seeds determines the results of K-means clustering. The conventional K-means algorithm usually adopts the random strategy to select initial seeds, which is unable to generate an ideal clustering result in many cases. To solve the problem of the existing initial seeds selection (abbreviated to ISS) strategies for K-means clustering, we propose a novel initial seeds selection algorithm, called ISS_OD, based on outlier detection. In ISS_OD, we select the initial seeds of K-means clustering by calculating the distance outlier factor of every object, the weighted density of every object and the weighted distances between objects. Experimental results on several UCI datasets demonstrate the effectiveness of our algorithm for the ISS of K-means clustering.

Read the paper · More papers on PaperTik