ChEsS: Cost-Effective Scheduling Across Multiple Heterogeneous Mapreduce Clusters

Nikos Zacheilas, Vana Kalogeraki · 2016

In recent years many organizations adopt the usage of multiple concurrent MapReduce frameworks running on different clusters in order to support data, failure, version and performance isolation for their Big Data applications. However, efficiently scheduling MapReduce workloads in such environments can be particularly challenging due to the observed tradeoff between the need for performance and the corresponding monetary cost. The problem is exacerbated by the fact that jobs have locality constraints and clusters employ different intrajob scheduling policies (e.g., FIFO, FAIR) for the execution of their jobs, affecting significantly the workload's execution time. In this paper we describe our approach for scheduling MapReduce jobs in multicluster environments taking into consideration the performance/budget tradeoff. Our approach makes the following contributions: (i) ChEsS, a novel Paretobased scheduling framework for identifying near-optimal jobsto-clusters assignments for user's workloads with respect to performance and cost, and (ii) a model that considers the impact of the different intra-job scheduling algorithms and the jobs' locality constraints on the observed performance and required budget. Our detailed experimental evaluation using both scientific and industry workload traces illustrate the working and benefits of our approach.

Read the paper · More papers on PaperTik