A common framework of partition-based clustering for large scale dataset using sampling and its MapReduce implementation
Ran Jin, Chunhai Kou, Ruijuan Liu, Guo, Tao · Tehnicki vjesnik - Technical Gazette · 2016
Original scientific paperClustering is one of the significant tasks in data mining, and partition-based clustering algorithms such as k-means are one of the popular solutions.However, with the increasing development of cloud computing and big data, large scale dataset has been a big challenge for clustering.For example, the execution of clustering algorithm is too time-consuming, the optimization of parameters is difficult, and the quality of clusters is not good.To this end, in this paper, we proposed a common framework of partition-based clustering algorithms such as k-means, and designed its MapReduce implementation.Specifically, in order to deal with the representation of large scale dataset, we propose to employ sampling technique.Then, inspired by k-means algorithm, we propose a common procedure of clustering, and provide a k-means based implementation.Furthermore, we implement proposed framework using MapReduce programming model.Experiments show that our method is efficient for large scale dataset.