Dynamic slot allocation technique for MapReduce clusters

Shanjiang Tang, Bu‐Sung Lee, Bingsheng He · 2013

MapReduce is a popular parallel computing paradigm for large-scale data processing in clusters and data centers. However, the slot utilization can be low, especially when Hadoop Fair Scheduler is used, due to the pre-allocation of slots among map and reduce tasks, and the order that map tasks followed by reduce tasks in a typical MapReduce environment. To address this problem, we propose to allow slots to be dynamically (re)allocated to either map or reduce tasks depending on their actual requirement. Specifically, we have proposed two types of Dynamic Hadoop Fair Scheduler (DHFS), for two different levels of fairness (i.e., cluster and pool level). The experimental results show that the proposed DHFS can improve the system performance significantly (by 32% ~ 55% for a single job and 44% ~ 68% for multiple jobs) while guaranteeing the fairness.

Read the paper · More papers on PaperTik