Towards adapting the mapreduce model for scientific applications

Madhusudhan Govindaraju, Zacharia Fadika · 2012

The suitability of MapReduce extends far beyond its web processing roots. The model can also be adapted for scientific applications. In this thesis we present specific optimizations suited for scientific applications and quantify the gains obtained. We study and test a variety of MapReduce applications, then subsequently identify the ones most suitable for our optimizations. In this thesis, we focus both on the performance of processing large-scale data in distributed scientific applications, as well as the processing of smaller but demanding input sizes primarily used in diskless, and memory resident systems. We present the design decisions and implementation tradeoffs for elastic frameworks that follow the MapReduce paradigm. We recognize the performance of current MapReduce implementations such as Hadoop and Twister to suffer in heterogeneous and load-imbalanced clusters. We present a list of recommendations for a solution capable of performing well not only in homogeneous settings, but also when clusters exhibit heterogeneous properties. We present the components and trade-offs necessary for efficient MapReduce implementations in heterogeneous cluster environments. Subsequently, we identify several MapReduce frameworks with various degrees of conformance to the key tenets of the model. Each however, optimized for specific features. HPC application and middleware developers must thus understand the dependencies between the specific features of a MapReduce framework and application requirements. We present a standard benchmark suite for quantifying, comparing, and contrasting the performance of MapReduce implementations under a wide range of representative use cases. We report the performance of three different MapReduce frameworks on the benchmarks, and draw conclusions about their current performance characteristics. The performance analysis we perform also throws light on the available design decisions for future implementations, and allows researchers to choose the MapReduce framework that best suits their applications' needs. We compare our approach in distributed environments over Apache Hadoop deployed in the Binghamton University Grid and Cloud Computing Laboratory, and also on the Magellan testbed at the National Energy Research Scientific Computing Center (NERSC). This thesis presents a set of design choices and related optimizations directly suitable for HPC environments.

Read the paper · More papers on PaperTik