Reversing statistics for scalable test databases generation
Entong Shen, Lyublena Antova · 2013
Testing the performance of database systems is commonly accomplished using synthetic data and workload generators such as TPC-H and TPC-DS. Customer data and workloads are hard to obtain due to their sensitive nature and prohibitively large sizes. As a result, oftentimes the data management systems are not properly tested before releasing, and performance-related bugs are commonly discovered after deployment, when the cost of fixing is very high. In this paper we propose RSGen, an approach to generating datasets out of customer metadata information, including schema, integrity constraints and statistics. RSGen enables generation of data that closely matches the customer environment, and is fast, scalable and extensible.