Skew‐aware online aggregation over joins through guided sampling

Yuxiang Wang, Jiahui Jin, Xiaoliang Xu, Longbin Zhang · Concurrency and Computation Practice and Experience · 2018

Summary Online aggregation is a query processing technique that returns approximate answers with error guarantees (in the form of confidence intervals) continuously during the query execution process. This approach offers users a suitable tradeoff between query efficiency and accuracy. The key issue of online aggregation is how to ensure a random sample collection's efficiency and effectiveness. However, the often‐used “blind” sampling method does not adequately consider dataset statistics and other useful information, leading to inefficient sampling and poor sample quality. This becomes a glaring performance issue for skewed data distribution over joins. To alleviate this problem, we utilize dataset statistics to propose a new “guided” sampling approach, which consists of a logic‐partition‐based weighted Gaussian sampling method tailored for the skewed join key, as well as a two‐level sample allocation method that applies to the skewed measured value. Extensive experiments using the TPC‐H benchmark for skewed data distribution demonstrate our solution's superior performance.

Read the paper · More papers on PaperTik