An Overview of Data Locality Based Scheduler in Hadoop MapReduce Environment
Khushboo Kalia, Neeraj Kumar Gupta · 2020
MapReduce is a processing framework used extensively for substantial set of data set. Default MapReduce schedulers are not efficient enough to process data sets of heterogeneous nature. Thereby, it leads to low performance and reduced data locality. Working in a heterogeneous environment require more memory, power and advance architecture and computation. This paper provides the outline on MapReduce scheduling, how it is affected by the data locality and heterogeneous data sets. Examination on data locality based scheduling mechanism is done, that aims in reducing the job's running time or CPU time and increases the efficiency, bandwidth and throughput of the cluster i.e. enhancing the overall performance of the Hadoop cluster. In addition to these merits, the lack point of this data locality based scheduling approach is also reviewed.