Improved Parallel Clustering with Optimal Initial Centroids

M.A. Anitha, K. A. Abdul Nazeer · 2017

Clustering of a large data set is one of the challenging tasks and it has much application in the areas such as bioinformatics, social networking, image segmentation and many others. k-means clustering is the most popular and widely used method in commercial applications and scientific research because of its simplicity. However, it has some disadvantages. The major issues are convergence to the local optima, that is, the quality of the clustering result is highly dependent on the initialization. Another problem is that the clustering will produce different results in different independent runs. The number of clusters have to be specified in advance. But in the real application, it is tough to determine the parameters in advance. This research is intended to develop a parallel clustering algorithm which is capable of clustering large data sets. An improved k-means type algorithm has been proposed that generate the optimal initial centroids using a new heuristic method and works on large data set using MapReduce methodology. The proposed method is accurate as compared to other existing methods of similar nature.

Read the paper · More papers on PaperTik