Efficient Data Clustering Algorithms: Improvements over Kmeans

Mohammed Abubaker, Wesam M. Ashour · International Journal of Intelligent Systems and Applications · 2013

This paper presents a new approach to overcome one of the most known disadvantages of the well-known Kmeans clustering algorith m.The problems of classical Kmeans are such as the problem of random init ialization of prototypes and the requirement of predefined number of clusters in the dataset.Randomly in itialized prototypes can often yield results to converge to local rather than global optimu m.A better result of Kmeans may be obtained by running it many times to get satisfactory results.The proposed algorith ms are based on a new novel definition of densities of data points which is based on the k-nearest neighbor method.By this definit ion we detect noise and outliers which affect Kmeans strongly, and obtained good initial prototypes from one run with automatic determination of K nu mber of clusters.This algorithm is referred to as Efficient In itializat ion of Kmeans (EI-Kmeans).Still Kmeans algorithm used to cluster data with convex shapes, similar sizes, and densities.Thus we develop a new clustering algorith m called Efficient Data Clustering Algorith m (EDCA) that uses our new definit ion of densities of data points.The results show that the proposed algorithms improve the data clustering by Kmeans.EDCA is able to detect clusters with different non-convex shapes, different sizes and densities.

Read the paper · More papers on PaperTik