A Density Based k-Means Initialization Scheme

Panagiotis Gourgaris, Christos H. Makris · 2015

In this paper we present the results from some versions of a new initialization scheme for the k-Means algorithm. k-Means is probably the most fundamental clustering algorithm, with application in lots of fields, such as Signal Processing, Image Colour Segmentation as well as Web data management. The initialization process of the algorithm is of great interest, raising two big challenges. The first one, is to find out what is k, the number of clusters. The second one, is to determine which are the initial k seeds. We here mainly focus on the later. Our approach is heuristic hence profound mathematical arguments are not being presented. We are based mainly on criteria like density, Euclidean distance and the Mardia's multivariate kurtosis statistic. In order to test the quality of our results, a few cluster validity measures, other than the commonly used Sum Squared Error(SSE) are applied, which in our belief are suitable to be used for evaluation purposes.

Read the paper · More papers on PaperTik