PUMA: Purdue MapReduce Benchmarks Suite

Faraz S. Ahmad, Seyong Lee, Mithuna S. Thottethodi, T. N. Vijaykumar · Purdue e-Pubs (Purdue University System) · 2012

MapReduce[5] is a well-known programming model, developed within Google, for processing large amounts of raw data such as crawled documents or web request logs on a cluster of commodity hardware comprising of thousands of machines.MapReduce provides automatic data management and fault tolerance to improve programmability of clusters.In the MapReduce programming model, programmers specify a Map function which processes input data to generate intermediate data in the form of tuples, and a Reduce function which further processes values associated with a particular key.Hadoop is an open-source implementation of MapReduce which is being improved and developed regularly by software developers / researchers and is maintained by Apache Software Foundation.Despite being vast efforts on the development of Hadoop MapReduce, there has not been a very rigorous work done on the benchmarks side.During our work on MaRCO[2], we developed a benchmark suite[8], called "PUMA" which represents a broad range of MapReduce applications exhibiting application characteristics with high/low computation and high/low shuffle volumes.There are a total of 13 benchmarks, out of which Tera-Sort, Word-Count, and Grep are from Hadoop distribution.The rest of the benchmarks were developed in-house and are currently not part of the Hadoop distribution.The three benchmarks from Hadoop distribution are also slightly modified to take number of reduce tasks as input from the user as well as to generate final time completion statistics of jobs. Benchmark detailsThe details of benchmarks, including command line execution format and input data set description can be found below.

Read the paper · More papers on PaperTik