Scalable storage managers for the multicore era
Anastasia Ailamaki, Franklin Johnson · 2010
Database Management Systems provide a crucial underpinning today's information-driven world, providing users with efficient and up-to-date access to huge volumes of data. In order to keep pace with exploding data volumes and increasingly sophisticated processing of that data, database engines must exploit fully the underlying hardware. Recent shifts in computer architecture have led to the rise of multicore designs which depend on parallelism peformance, with the result that hardware advances no longer deliver performance for free. Today's software must be scalable , providing exponentially-increasing parallelism to keep the underlying hardware busy. While database workloads have the advantage of high concurrency both within and between requests, database engines have historically focused on using that concurrency to overlap delays more than to exploit parallelism heavily within a single machine. As a result, bottlenecks internal to the database engine itself prevent it from converting concurrency into sufficient parallelism, especially in transaction processing workloads. In this thesis, we show how to move the database engine off the critical path, proving that database engines can achieve the scalability needed to exploit today's parallel hardware. We identify three key areas achieving this goal. First, performance optimizations must focus first on the critical path. Every serial computation is a liability, while single-thread performance and even the total amount of work performed by a computation are secondary concerns. Second, in order to remain scalable as hardware parallelism continues to increase, bottlenecks must be eliminated—not just reduced—by removing the source of the contention. Finally, we identify scheduling as a critical area current and future systems. Improper scheduling can increase artificially the length of the critical path in the system, while effective scheduling can eliminate many bottlenecks by changing access patterns and improving regularity in the system. We demonstrate the effectiveness of the above approaches by applying them to state-of-the-art database systems running on highly parallel multicore hardware. These results also generalize to the wider software community, as concerns with critical path, bottlenecks, and scheduling arise in every software domain. Finally, this work demonstrates that many of the remaining challenges in achieving database engine scalability lie with scheduling, suggesting a path toward scalability-enhancing scheduling techniques.