Improving MapReduce Load Balancing in Hadoop

Fu-Hong Syue, Varsha A. Kshirsagar, Shou‐Chih Lo · 2018

MapReduce is a parallel processing framework widely used in big data processing in recent years with strengths such as simple operation, high fault tolerance, and high scalability. Data skew is a typical problem when MapReduce is used to handle data-intensive applications. If the default key partitioning function of Hadoop is applied to process such datasets, the workload will not be evenly distributed to each reducer under most circumstances. In this paper, we propose a load balancing mechanism to mitigate the negative effect of data skew on the performance of MapReduce. Proposed mechanisms combine the reservoir sampling and greedy algorithms, and further incorporate the concept of data locality. The experimental results demonstrate that the proposed load balancing mechanism can effectively reduce the data transmission cost through the network to each reducer. More importantly, this mechanism can balance the workload of each reducer.

Read the paper · More papers on PaperTik