DataSynth

Arvind Arasu, Raghav Kaushik, Jian Li · Proceedings of the VLDB Endowment · 2011

A variety of scenarios such as database system and application testing, data masking, and benchmarking require synthetic database instances, often having complex data characteristics. We present DataSynth , a flexible tool for generating synthetic databases. DataSynth uses a simple and powerful declarative abstraction based on cardinality constraints to specify data characteristics, and uses sophisticated algorithms to efficiently generate database instances satisfying the specified characteristics. The demo will showcase various features of DataSynth using two real-world data generation scenarios.

Read the paper · More papers on PaperTik