MR-KClust: An efficient Map Reduce based clustering Technique

Snigdha Agarwal, Rituparna Sinha · 2022 6th International Conference on Electronics, Communication and Aerospace Technology · 2022

The world is rich in data, but efficient storage and processing of the data are of the utmost importance. As the data grows, processing it and applying algorithms to extract meaningful and hidden patterns efficiently becomes a challenge for the research community. Traditional clustering approaches, such as K Means and its variations, may not perform well when working with big data, that is, if the input data is large and has a high computational cost and slow speed Keeping this in mind, an efficient map-reduced clustering algorithm, named MR-KClust, has been designed It is a variant of the well-known k-means clustering, where a polynomial regression has been applied for the initial selection of centroids and algorithms. In the Map Reduce design of the algorithm, the data has been split, and each input split is provided to the mappers, where the map function calculates the distances from each sample to each centroid point and emits relabeled cluster ID as a key and data ID coordinates as a value. The combiner function combines samples with the same key for each machine and loads what the mapper emits. The performance of the algorithm as compared with the existing variants of K - Means outperformed the others with respect to the mean squared error and has a decent runtime.

Read the paper · More papers on PaperTik