Execution Time Prediction for Apache Spark

Zhipeng Gao, Ting Wang, Qian Wang, Yang Yang · 2018

Apache Spark is a framework that being increasingly used in distributed data processing. However, the performance of a Spark application can vary considerably depending on many factors, including the input data, implementation of the program, Spark configuration parameters and cluster resources, making performance prediction becomes an arduous work. To address this challenge, in this paper, we design a two-steps prediction framework with high prediction accuracy. The technical core is that each cluster of applications with similar performance behavior has a particular performance prediction model applying gradient boosting regression. The performance behavior is profiled using both CPU and memory utilization series. It adapts to ad-hoc applications by simulating the execution using limited amount of sample input data, matching the resource utilization signature with history applications, and predicting execution time using corresponding performance model. Experiment evaluates that the framework can predict execution time for ad-hoc applications with high accuracy and efficiency.

Read the paper · More papers on PaperTik