Data processing on heterogeneous hardware

Max Heimel · DepositOnce · 2018

The primary objective of data processing research on modern hardware is to understand how to utilize emerging technology to process data efficiently. Over the last decades, Software Engineers and Computer Scientists have made significant progress towards this goal, providing highly-tuned algorithms, systems & mechanisms for a wide variety of different device types. However, while we mostly understand how to exploit multiple identical devices, the question of how to build database engines that can deal with a diverse number of them remains open. Already today, with multi-core CPUs, graphics cards, FPGAs, and vector processors, we have access to a staggering variety of different architectures, and experts agree that hardware diversity will continue to grow. As we enter this era of heterogeneity, developing new approaches to deal with a multitude of different hardware architectures is becoming increasingly important. Tomorrow’s database systems will need to exploit and embrace this trend towards increased hardware heterogeneity to meet the performance requirements of the modern information society. In this thesis, we present our contributions to the area of data processing on heterogeneous hardware. In particular, we discuss the following two major topics: In the first part of this thesis, we discuss how to manage hardware heterogeneity by reducing the development overhead to implement a database engine that can run on different hardware architectures. For this, we propose an alternative system design, based on a single set of hardware-oblivious operators which are compiled down to the actual hardware at runtime. Relying on hardware abstraction, learning mechanisms and self-tuning algorithms, this approach can drastically reduce the enormous development overhead that typically comes with supporting a variety of architectures, while still achieving competitive performance to hand-tuned code. We also present Ocelot, a prototypical hardware-oblivious database engine to demonstrate the feasibility of this concept. In the second part of this thesis, we introduce a novel approach to exploit hardware heterogeneity in a relational database engine by utilizing graphics processing units to assist the query optimizer. Based on this idea of GPU-Assisted Query Optimization, we then introduce a self-tuning, multivariate selectivity estimator based on Kernel Density Estimation that is specifically designed to exploit the properties of modern graphics cards and multi-core CPUs. This approach enables us to develop a novel estimator that is highly scalable, can adapt itself to changes in both the database and the query workload, and typically outperforms the accuracy of established state-of-the-art methods. For both topics, we motivate our decisions, present and discuss experimental evaluations, provide both the source code and scripts to allow other authors to reproduce our results, and outline potential directions for future work.

Read the paper · More papers on PaperTik