An evaluation of MapReduce framework in cluster analysis
Tanvir Habib Sardar, Ahmed Rimaz Faizabadi, Zahid Ansari · 2017
Data is growing exponentially due to the World Wide Web and scientific enhancements, requires proper strategies and techniques to deal with it. Thus, it requires increasing the computational requirements. Efficient processing paradigms and smart implementation architecture is the key to meet the scalability and performance requirements in large scale data analysis. Clustering is one of the popular data mining techniques to analyse datasets. MapReduce programming paradigm, on the top of Hadoop distributed architecture, is widely used in today's world in order to obtain efficiency for large dataset clustering. In this work, we have experimented and validated the efficiency of K-means clustering algorithm in MapReduce paradigm over hadoop architecture while processing different sized datasets in combination of different hadoop cluster sizes.