Pareto-Based Scheduling of MapReduce Workloads

Nikos Zacheilas, Vana Kalogeraki · 2016

In recent years we are observing an increased demand for processing large amounts of data. The MapReduce programming model has been utilized by major computing companies in order to perform large-scale data processing. However, the problem of efficiently scheduling MapReduce workloads in cluster environments, like Amazon's EC2, can be challenging due to the observed tradeoff between the need for performance and the corresponding monetary cost. The problem is exacerbated by the fact that cloud providers tend to charge users based on their I/O operations increasing dramatically the spending budget. In this paper we describe our approach for scheduling MapReduce workloads in cluster environments taking into consideration the performance/budget tradeoff. Our approach makes the following contributions: (i) a novel Pareto-based scheduler for identifying near-optimal resource allocations for user's workloads with respect to performance and monetary cost, and (ii) automatic configuration of tasks' buffer sizes to minimize the I/Os impact on the users' budget. Our detailed experimental evaluation using both real and synthetic datasets illustrate that our approach can improve the performance of the workloads as much as 50%, compared to its competitors.

Read the paper · More papers on PaperTik