ACORE: A Query Optimization Approach for Spark SQL Based on Cost Model and Markov Prediction Model

Lanxin Su, Huiyong Liu · Journal of Physics Conference Series · 2021

Abstract Spark SQL is Spark’s module for working with structured data. However, there exist two problems that may impact its query execution efficiency: 1) Users may attempt to cache datasets used multiple times in a query to speed up the execution, but due to the predicate pushdown optimization and the overhead of writing cached data, the cached plan may turn out to have worse performance. 2) When using Spark dynamic resource allocation strategy, the release of idle executors would lead to the loss of cached data, introducing recalculation cost. To address these, we present an approach, named ACORE, which adaptively cache datasets and optionally release executors to optimize query execution of Spark SQL. The approach includes two models: a cost model and a Markov chain prediction model. The cost model is applied to estimate the execution cost for cached and un-cached plans, and decide which one to be the actual execution plan. The Markov prediction model is used to predict the changing trend of intervals between queries, guiding to optionally release executors to save cached data. We test the efficiency of our approach with data generated by TPC-H tool and the experimental results show that the execution performance can be improved by up to 51%.

Read the paper · More papers on PaperTik