Phase Annotated Learning for Apache Spark: Workload Recognition and Characterization

Seyedali Jokar Jandaghi, Arnamoy Bhattacharyya, Cristiana Amza · 2018

In this paper, we introduce and evaluate a novel resource modeling technique for workload profiling, detection and resource usage prediction for Spark workloads. Specifically, we profile and annotate resource usage data in Spark with the application contexts where the resources were used. We then model the resource usage, per context, based on a Mixture of Gaussians (MOG) probabilistic distribution technique. When we recognize a similar workload, we can thus predict its resource usage for the contexts modeled a priori. In order to experimentally test the functionality of our Spark stage annotator and workload modeling tool we performed workload profiling for eight Apache Spark workloads. Our results show that, whenever a previously modeled workload is detected, our MOG models can be used to predict resource consumption with high accuracy.

Read the paper · More papers on PaperTik