A Machine Learning Approach for Predicting Execution Statistics of Spark Application
Piyush Sewal, Hari Singh · 2022
Apache Spark is one of the most popular, widely used and open-source distributed processing framework that can process huge site datasets in time efficient manner due to its in-memory computational capabilities. However, there are several factors that can affect the performance of an application which include the nature and size of the input dataset, computational capability of the system and nature and design of the algorithm. Hence, there are different parameters that are required to correctly predict the execution statistics of a Spark application which include execution time of jobs, stages and tasks, memory requirement and usage at the execution level and I/O cost in the form of read/ write shuffling of data. To address these challenges, a simulation and machine learning based prediction model is presented in this paper that takes only a few initial samples of execution statistics and predicts the performance and execution statistics of the Spark application with high accuracy. The proposed model is evaluated on the Wordcount application and Spark standalone mode and accuracy metrics show that the proposed model achieves high accuracy in predicting execution statistics.