Benchmarking Big Data SQL Frameworks

Vlad Mirel · Aaltodoc (Aalto University) · 2016

The amount of data being generated on a daily basis is constantly increasing, pushing the limits of traditional data processing technologies.A consequence of this increase is the rise to new distributed Big Data engines.This thesis is focused on benchmarking Big Data SQL frameworks, both open-source or proprietary.The Big Data frameworks are compared with each other from three points of view: performance (total job execution time), feature availability and integration with other services.In order to provide an unbiased comparison, a similar underlying infrastructure was employed for each framework.More precisely, experiments were conducted on different Big Data SQL platforms hosted on two public cloud infrastructures: Microsoft Azure and Google Cloud Platform.In the case of Azure, SQL queries were executed on HDInsight, a PaaS solution for Big Data SQL clusters like Spark SQL, HiveQL, Apache Drill and Apache Impala.Experiments were also conducted on SaaS solutions offered by the both vendors, Microsoft Azure Data Lake Analytics and Google BigQuery.The workloads comprised from several GBs up to 250 GBs in Parquet format.In the case of SaaS platforms, 44.8 GBs of .csvfiles were employed.The results obtained from conducting the experiments on both PaaS and SaaS platforms are meant to shed some light on the benefits that emerge when choosing one technology.Furthermore, based on these insights, existing Big Data engines could be further improved.

Read the paper · More papers on PaperTik