Improved k-means Algorithm using Density Estimation
Md Abdul Masud, Md. Moshiur Rahman, Shaneworn Bhadra, Subrata Saha · 2019
Unsupervised learning approach, clustering, performs a set of groups where similar types of data points are placed in a group. K-means is a popular clustering algorithm, which is commonly used in different application because of its simplicity and efficiency. This main drawback of k-means algorithm is random selection of initial cluster centers. This paper proposes an improved k-means, namely, IK-means algorithm to overcome the main pitfall of k-means. The IK-means algorithm uses Kd-tree data structure to present and store data objects, and applies kernel density estimation technique to locate the densest areas of data points. Initial cluster centers are assigned from the densest areas. The IK-means algorithm produces the clustering results with better accuracy that improves the performance of the k-means algorithm on artificially generated and real datasets.