An Efficient MapReduce-Based Adaptive K-Means Clustering for Large Dataset
Tapan Chowdhury, Arijit Mukherjee, Susanta Kumar Chakraborty · 2017
MapReduce-based clustering of a large dataset is gaining rapid importance in the fields of data science. Its main task is to make groups of data from the large dataset such that all data points categorized into a single group are similar to one another. K-means clustering method is one of the most extensively used clustering methods but it minimizes clustering criteria by iteratively relocating data points between clusters until a locally optimal partition is attained, so convergence is local and globally optimal solutions cannot be guaranteed in case of a large dataset. Due to the random selection of K-initial seeds, it decreases the quality of clusters. This paper proposes a MapReduce-based adaptive K-means clustering approach. The adaptive nature of our approach, improves the efficiency based on two key concepts. First, it selects the initial seeds that are spread throughout the large dataset via statistical testing. Second, it reduces the impact of the outlier on a large dataset and clustering large dataset using MapReduce framework. Our approach is adaptive for large dataset by adapting the size and nature of the dataset to make the process more robust and efficient. The experimental result shows that our approach improves the performance of clustering compared to earlier works.