A method of classification-based Spark job performance modeling
zhiyong Ding, Chaohui Zhang · 2nd International Conference on Applied Mathematics, Modelling, and Intelligent Computing (CAMMIC 2022) · 2022
Prediction of Apache Spark job execution time is a key technology to guide Spark cluster resource allocation and parameter tuning. In the existing research, a unified modeling method is used for different jobs, and the prediction model considers less factors, resulting in poor prediction effect. In view of the above problems, this paper proposes a classification-based Spark job performance modeling method. The method first selects features that are strongly correlated with job execution time, then classifies jobs according to the selected features, and finally uses GBDT algorithm to build an execution time prediction model for each class of jobs classified. The experimental results show that, compared with the method using unified modeling, the method proposed in this paper can reduce the RMSE and MAPE of the prediction results by an average of 42.5% and 51.1%.