ApproxJoin

Do Le Quoc, İstemi Ekin Akkuş, Pramod Bhatotia, Spyros Blanas, Ruichuan Chen, Christof W. Fetzer, Thorsten Strufe · 2018

A distributed join is a fundamental operation for processing massive datasets in parallel. Unfortunately, computing an equi-join over such datasets is very resource-intensive, even when done in parallel. Given this cost, the equi-join operator becomes a natural candidate for optimization using approximation techniques, which allow users to trade accuracy for latency. Finding the right approximation technique for joins, however, is a challenging task. Sampling, in particular, cannot be directly used in joins; naïvely performing a join over a sample of the dataset will not preserve statistical properties of the query result.

Read the paper · More papers on PaperTik