Accelerating array joining with integrated value-index
Haoyuan Xing, Gagan Agrawal · 2019
Large-scale multidimensional array processing is becoming an increasingly important problem with the rise of big data, scientific data processing, and machine learning workloads. One of the prevalent query types, array join, compares the cells of two arrays and finds matching cell pairs and is useful in finding patterns and differences across multiple arrays. These queries of ten have two distinct characteristics: first, the queries usually have a value-based filter predicate, because one may only be interested in a small subset of the entire array, and second, because many multidimensional data are numerical and has inaccuracies in nature, the comparison between cells are of ten required to be approximate. These two characteristics motivate the need for value similarity array joins. While prior works exist in both array join and similarity join, they primarily target dimension-based comparisons and do not address this problem.