A novel K-means based clustering algorithm for big data

Ankita Sinha, Prasanta K. Jana · 2016

Data generation has seen tremendous growth in the past decade. Managing such huge amount of data is a big challenge. Clustering can serve as a solution, it divides the data into smaller groups based on the level of similarity among the objects. K-Means is one of the most popular and robust clustering algorithm. However, the major drawback of K-Means is to input the number of clusters which is not known in advance particularly for real world data sets. In this paper, we propose a K-Means based clustering algorithm for big data in which we automate the number of clusters to deal with big data. The algorithm is implemented using Spark, a better programming framework than the MapReduce. The proposed algorithm is simulated extensively with large scale synthetic data set as well as real life data on a 4 node cluster. The simulated results demonstrate better performance of the proposed algorithm over the scalable K-Means++ implemented in MLLIB library of Spark.

Read the paper · More papers on PaperTik