Partitioning and Bucketing Techniques to Speed up Query Processing in Spark-SQL

Yassine Ramdane, Omar Boussaïd, Nadia Kabachi, Fadila Bentayeb · 2018

Horizontal partitioning is an optimization technique applied to improve query processing time in distributed data warehouses. Scanning a large number of HDFS data blocks, to respond to the ad-hoc or OLAP queries, is a heavy operation. We can skip loading unnecessary data blocks if we partition or index some tables by the appropriate predicate attributes. However, the way of selecting the candidate's attributes remains a challenging task to handle. In this paper, we propose a technique based on frequent itemset mining, to Partition, Bucket and Sort the Tables (PBSTs) of a big data warehouse with the more frequent predicate attributes in the queries. We take into account the density of the attributes of the tables, data skew, and the physical characteristics of the cluster nodes. To evaluate our approach, we conducted some experiments in a cluster of 15 slave nodes. Experimental results show that with our method, we improve the query response time by 50 % over existing skipping techniques.

Read the paper · More papers on PaperTik