h-MapReduce: A Framework for Workload Balancing in MapReduce

VenkataSwamy Martha, Weizhong Zhao, Xiaowei Xu · 2013

The big data analytics community has accepted MapReduce as a programming model for processing massive data on distributed systems such as a Hadoop cluster. MapReduce has been evolving to improve its performance. We identified skewed workload among workers in the MapReduce ecosystem. The problem of skewed workload is of serious concern for massive data processing. We tackled the workload balancing issue by introducing a hierarchical MapReduce, or h-MapReduce for short. h-MapReduce identifies a heavy task by a properly defined cost function. The heavy task is divided into child tasks that are distributed among available workers as a new job in MapReduce framework. The invocation of new jobs from a task poses several challenges that are addressed by h-MapReduce. Our experiments on h-MapReduce proved the performance gain over standard MapReduce for data-intensive algorithms. More specifically, the increase of the performance gain is exponential in terms of the size of the networks. In addition to the exponential performance gains, our investigations also found a negative effect of deploying h-MapReduce due to an inappropriate definition of heavy tasks, which provides us a guideline for an effective application of h-MapReduce.

Read the paper · More papers on PaperTik