SARAH - Statistical Analysis for Resource Allocation in Hadoop
Bruce Carruthers Martin · 2014
Improving the performance of big data applications requires understanding the size and distribution of the input and intermediate data sets. Obtaining this understanding and then translating it into resource settings is challenging. SARAH provides a set of tools that analyze input and intermediate data sets and recommend configuration settings and performance optimizations. Statistics generated by SARAH are persistently stored, incrementally updated and operate across the several processing frameworks available in Apache Hadoop. In this paper we present the SARAH tool set, describe several Hadoop use cases for utilizing statistics and illustrate the effectiveness of utilizing statistics for balancing reduce workload on Map-Reduce jobs on web server log file data.