Comparative Benchmarking of Relational Database Systems
Carolyn Turbyfill · 1988
New and diverse architectures have been developed to meet rising expectations of both functionality and performance from relational database systems. A comprehensive, portable, comparative benchmark is needed to evaluate the performance tradeoffs inherent in these different architectures. In this thesis, we develop an experimental framework for systematically evaluating and comparing relational database systems across diverse architectures. We describe and motivate the design of a scalable, portable benchmark for relational database systems, the AS$\sp3$AP benchmark (ANSI SQL Standard Scalable And Portable). The AS$\sp3$AP benchmark is designed to provide meaningful measures of database processing power and to be a useful tool for system designers. We introduce a new performance metric, the equivalent database ratio, to be used in comparing systems. The equivalent database ratio is the ratio of database sizes on two different systems for which equivalent performance on a set of test queries is obtained. The benchmark consists of two parts: tests of the access methods and building blocks, and test of the optimizer. In tests of the access methods and building blocks, access methods and functions common to implementations of relational database systems are tested. The relations and queries used in these tests are designed to increase the likelihood that the access method or program branch targeted by the test will actually be chosen by the query optimizer. The tests of the access methods and building blocks are used to compute the equivalent database ratio. In tests of the optimizer, assumptions typically made by query optimizers such as the uniform distribution of data values or the random placement of tuples are systematically violated. Comparisons between queries are used to indicate whether the optimizer has made a correct choice. Queries are included to illustrate current limitations in the state of the art in query optimization. For a representative cross section of access methods, estimates of the best case response time for the benchmark queries on disk bound systems are computed. The response time estimates and the systematic benchmarking methodology provide for clear interpretation of benchmark results.