Fine tuning of MapReduce jobs using parallel K Map clustering
Divya Mishra, Suryakant yadav · Journal of Emerging Technologies and Innovative Research · 2019
Advancement and development of information technology has led to huge growth in data, which pose huge challenge into data storage and analysis to conclude meaningful information from data. Size of data ranges from petabyte or exabyte range, to mine information out of large data set will increase the computing capacity. Thus to address and manage this high velocity data growth an advanced processing algorithms and methods are required for data analysis. To achieve same MapReduce word count program and P-KMeans a parallel clustering algorithm used in Hadoop .Through experiment, it has been concluded that execution time reduce when nodes count increases in a cluster, but also some of the important points has been observed while conducting experiment. Various performance change, and plotted results on different performance charts. This paper aims to study MapReduce applications and its performance verification and improvement recorded for Parallel KMeans algorithm on four nodes Hadoop Cluster.