Locality premised reducer scheduling in Hadoop

Nusrat Fatma, Hemant Kr. Singh, Shafeeq Ahmad, Prachi Srivastava · 2016

Data is generated in high speed and volume called Big Data. Hadoop has ability to handle Big Data. Programs get executed in parallel manner in Hadoop. In Hadoop the problem gets divide in two parts namely, map task and reduce task. In the existing Hadoop version map task scheduling premise is the locality of input data lowers the network traffic and hence improve performance of the mappers. But reduce task get scheduled without any consideration of data locality, resulting to poor performance at requesting node. This paper propose a modified reduce task scheduling algorithm on the basis of data locality that will minimize data-local traffic. In the evaluation of algorithm it is observed that up to 80 % of bytes shuffling has reduced in Hadoop clusters.

Read the paper · More papers on PaperTik