Data locality in Hadoop cluster systems
Mukhtaj Khan, Yang Liu, Maozhen Li · 2014
MapReduce has become a major programming model that supports distributed and parallel processing for large-scale data-intensive applications such as Web data mining, network traffic analysis, machine learning and scientific simulation. Hadoop is the most popular open-source implementation of the MapReduce programming model. In Hadoop, input files are divided into many data blocks and these blocks are distributed over several nodes in cluster. To efficiently process the data blocks, Hadoop should provide an efficient scheduling mechanism for enhancing the performance of the system in a shared cluster environment. In Hadoop scheduling mainly caused by data locality issues due to limited network bandwidth. By introducing the scheduling issues with regarding to the data locality, this paper review different data locality aware scheduling algorithms that handling the data locality issues. In addition, this paper also evaluating their features, strength, weakness and provided some guidelines on how to improve further these scheduling algorithms.