Parallel SQL Query Auto-Tuning on Multicore
Victor Pankratius, Martin Heneka · Repository KITopen (Karlsruhe Institute of Technology) · 2011
Multicore processors with several processors on a chip are standard, so applications need to be parallel in order to exploit the performance potential. Relational database systems are important applications that can exploit new opportunities for parallelism within queries. Intra-query parallelism offers additional performance potential that could not be exploited easily on earlier hardware. Addressing this important issue, this paper focuses on a difficult scenario for performance improvement: the parallelization of joins in I/O intensive multi-join queries. Our approach has the significant advantage that is does not require a rewrite of existing query optimizers from scratch. We boost query speed on multicore systems using query execution plans that are generated by sequential optimizers. This is the first paper to (1) auto-tune parallel query performance by adjusting the structure of multi-threaded pipelines that are superimposed over sequential query plans; (2) employ double-pipelined hash joins that are multithreaded to boost performance; (3) let an auto-tuner decide how to adapt parallelism to the hardware environment by exchanging hash join algorithms in the query plan; (4) present a demonstration and working strategies for multithreaded query auto-tuning on shared-memory multicore systems. Our evaluations show that queries can execute up to a factor of 3 faster on a quad-core machine, and that queries from the industry TPC-H benchmark can execute up to 47% faster compared to sequential execution. The results are remarkable considering the I/O bound context. Using PostgreSQL’s code as an example, we also discuss the software engineering issues for the adaptation of real-world database systems.