A Collaborative Framework for Similarity Enforcement in Synthetic Scaling of Relational Datasets
Jiangwei W. Zhang, Yong Chiang Tay · 2019
Researchers and developers use benchmarks to compare their algorithms and products. A database benchmark must have a dataset. To be application-specific, this dataset should be empirical. However, the dataset may be too small, or too large, for the benchmarking experiments. The dataset must, therefore, be scaled to the desired size. To ensure the scaled dataset is similar to the original dataset, previous work typically specifies or extracts a fixed set of properties from the original dataset, then uses these properties to generate synthetic data for the scaled dataset. However, this approach becomes increasingly intractable as the property set gets larger, so a new solution is necessary. This paper proposes ASPECT, which adopts a different approach, for relational datasets. The user first selects a sizescaler to synthetically scale the empirical dataset to a desired size, then uses a tool to tweak the scaled dataset to enforce each target property. ASPECT provides the interface for these tools to collaborate in this tweaking process. Extensive experiments on real datasets show that ASPECT can efficiently and effectively enforce similarity in such a generated dataset.