Improving Fair Scheduling Performance on Hadoop

Ya-Wen Cheng, Shou‐Chih Lo · 2017

Cloud computing is a potential technique to deal with big data. Apache Hadoop which provides the MapReduce parallel processing framework becomes a popular system for distributed storage and processing of large data sets on computer clusters. The performance of Hadoop in parallel data processing is relied on the efficiency of a MapReduce scheduling algorithm underlying. In this paper, we improve the performance of the well-known fair scheduling algorithm adopted in Hadoop by introducing several mechanisms. The modified scheduling algorithm can properly adapt to the runtime environment's condition with the objective of job fairness and short response time. Performance evaluations verify the superiority of the proposed algorithm over the original fair sharing algorithm.

Read the paper · More papers on PaperTik