Automatic Tuning of SQL-on-Hadoop Engines on Cloud Platforms

Prasad Deshpande, Amogh Margoor, Rajat Venkatesh · 2018

More and more companies are running Big Data workloads on cloud platforms. Configuration tuning of these data engines continues to be an essential but difficult undertaking. Cloud platforms further add to this complexity due to elasticity of compute resources and availability of different machine types. Data engineers have to choose the machine type and the cluster size in addition to the other configuration options. In this paper, we address the problem of automatically determining good configuration parameters for SQL workloads on SQL-On-Hadoop Engines like Hive and Spark, with the goal of optimizing resource usage and cost. We chose to focus on SQL workloads because SQL or SQL-like languages continue to be the most popular choice of ETL engineers and analysts. We propose a model based approach that relies on simple rules and insights into SQL Data Engine behavior and evaluate its effectiveness on both synthetic and real workloads. We show that very simple models can provide very good recommendations compared to the default configuration. In fact, we found that these were often better than those chosen by experts on real customer workloads. A somewhat surprising result is that more expensive machines with larger RAMs often do not lead to significant improvements in query execution times. The principles behind our model based approach are generic and can be adopted for any big data engine.

Read the paper · More papers on PaperTik