The Performance Characteristics of MapReduce Applications on Scalable Clusters
Kenneth Wottrich · 2011
Many cluster owners and operators have begun to consider MapReduce as a viable processing paradigm for its ability to address immense quantities of data. However, the performance characteristics of MapReduce applications on large clusters are generally unknown. The purpose of our research is to the characterize and model the performance of MapReduce applications on typical, scalable clusters based on fundamental application data and processing metrics. We identied ve fundamental characteristics which dene the performance of MapReduce applications. We then created ve separate benchmark tests, each designed to isolate and test a single characteristic. The results of these benchmarks are helpful in constructing a model for MapReduce applications.