Exploiting Cloud Heterogeneity to Optimize Performance and Cost of MapReduce Processing

Zhuoyao Zhang, Ludmila A. Cherkasova, Boon Thau Loo · ACM SIGMETRICS Performance Evaluation Review · 2015

Cloud computing offers a new, attractive option to customers for quickly provisioning any size Hadoop cluster, consuming resources as a service, executing their MapReduce workload, and then paying for the time these resources were used. One of the open questions in such environments is the right choice of resources (and their amount) a user should lease from the service provider. Typically, there is a variety of different types of VM instances in the Cloud (e.g., small, medium, or large EC2 instances). The capacity differences of the offered VMs are reflected in VM's pricing. Therefore, for the same price a user can get a variety of Hadoop clusters based on different VM instance types. We observe that the performance of MapReduce applications may vary significantly on different platforms. This makes a selection of the best cost/performance platform for a given workload a non-trivial problem, especially when it contains multiple jobs with different platform preferences. We aim to solve the following problem: given a completion time target for a set of MapReduce jobs, determine a homogeneous or heterogeneous Hadoop cluster configuration (i.e., the number, types of VMs, and the job schedule) for processing these jobs within a given deadline while minimizing the rented infrastructure cost. In this work,1 we design an efficient and fast simulation-based framework for evaluating and selecting the right underlying platform for achieving the desirable Service Level Objectives (SLOs). Our evaluation study with Amazon EC2 platform reveals that for different workload mixes, an optimized platform choice may result in 45-68% cost savings for achieving the same performance objectives when using different (but seemingly equivalent) choices. Moreover, depending on a workload the heterogeneous solution may outperform the homogeneous cluster solution by 26-42%. We provide additional insights explaining the obtained results by profiling the performance characteristics of used applications and underlying EC2 platforms. The results of our simulation study are validated through experiments with Hadoop clusters deployed on different Amazon EC2 instances.

Read the paper · More papers on PaperTik