A Collaborative Grouping Aggregation Query Scheme on Heterogeneous Computing Systems
Kailai Zhou, Xinwei Feng · 2022
Aggregation is one of the most common operations in data analytics systems, but complex aggregation query like grouping aggregation on massive data sets is a relatively expensive operation. Thus, improving the efficiency of aggregation query has become a hot topic for database researchers. The emergence of massively parallel coprocessors provides a way to solve this issue. Prior works focus mainly on aggregation optimization using one kind of coprocessors, but little is known about the performance impact on the heterogeneous computing system with diverse coprocessors, such as GPGPUs, MICs or FPGAs. To investigate the potential benefits of the use of heterogeneous computing systems in improving the performance of aggregation operation, in this paper, we design a grouping aggregation scheme with GPUs and MICs to work collaboratively. The scheme includes task partitioning method, data transfer mode, and aggregation algorithm optimization. To verify the effectiveness of our collective aggregation scheme, we performed a series of experiments on a heterogeneous platform with one Intel Xeon Phi coprocessor and one NVIDIA GPU coprocessor, and the results have shown that our aggregation processing strategy for the heterogeneous systems is effective and can achieve noticeable performance improvement comparing the CPU-only parallel computing system, even taking the communication overhead via PCI Express bus into account.