Impact of Map-Reduce framework on Hadoop and Spark MR Application Performance
Ishaan Lagwankar, Ananth Narayan Sankaranarayanan, Subramaniam Kalambur · 2020
Hadoop and Spark are popular Map Reduce frameworks, and are under active maintenance. Various benchmark suites and micro-benchmarks have been developed in order to measure and understand the performance and behaviour of the framework. The performance and behaviour of the micro-benchmarks (such as data motifs) are derived from core computations of Big Data applications can also be better understood by such an analysis. Data motifs have been previously studied and each data motif has been shown to behave differently due to changes in two external factors: its input size or input pattern. However, what is not considered by these studies or by any other Hadoop/Spark micro-benchmark distribution, is how the underlying framework can impact the behaviour of the benchmark, in terms of system resource usage and in terms of micro-architectural behaviour. The framework is the third external factor which impacts the behaviour of a big data application. In this work, we analyse the various Hadoop and Spark micro-benchmarks with different input sizes, and differing Hadoop/Spark versions and demonstrate through our results that the behaviour of the motif must also include the underlying framework as a third influencing factor.