Using Realistic Simulation to Identify I/O Bottlenecks in MapReduce Setups
Guanying Wang, Ali Raza Butt, Prashant Pandey, Karan Gupta · 2008
The exponentially growing data demands of modern enterprise and scientific applications poses critical challenges in sustaining the applications at scale. The MapReduce [1] programming model has served as the key enabler for executing resource-intensive applications over huge datasets. However, its configuration design-space has not been studied in detail. This is a complex problem as a typical MapReduce configuration can encompass hundreds of parameters, e.g., node configuration (number of disks and compute capacity), network topology (inter and intra-rack), choice of file system, data partitioning and layout, types of schedulers, etc – all of which affect application performance. While empirical insights for certain specific configurations, e.g., Google’s MapReduce infrastructure [1], do exist, they cannot be simply extended to other setups. Moreover, no tool or model is available to the community for studyingMapReduce application performance. In this work, we explore how choices about cluster design, run-time parameters, multi-tenancy and application design, affect I/O patterns, network communication and performance of MapReduce applications. Since the scale of the system precludes using actual machines for this exploration, we are developing an accurate MapReduce simulator, Dumbo, to facilitate performance analysis. The insights gained through Dumbo will be useful in comprehending the factors that affect MapReduce application performance. We expect Dumbo to be used by researchers and practitioners to understand how their MapReduce applications will behave on a particular configuration, and how they can improve the applications and platforms to optimize performance. Dumbo, used as a planning tool, will make MapReduce deployment far easier by reducing the number of parameters that currently have to be hand-tuned using trial-and-error and rules of thumb.