Cross-engine query execution in federated database systems
Ankush M. Gupta, Vijay N. Gadepally, Michael R Stonebraker · 2016
We have developed a reference implementation of the BigDAWG system: a new architecture for future Big Data applications, guided by the philosophy that “one size does not fit all”. Such applications not only call for large-scale analytics, but also for real-time streaming support, smaller analytics at interactive speeds, data visualization, and cross-storage-system queries. The importance and effectiveness of such a system has been demonstrated in a hospital application using data from an intensive care unit (ICU). In this article, we describe the implementation and evaluation of the cross-system Query Executor. In particular, we focus on cross-engine shuffle joins within the BigDAWG system, and evaluate various strategies of computing them when faced with varying degrees of data skew.