HYBRID SYNTHETIC DATA GENERATION FOR RELATIONAL DATABASES USING CONSTRAINT-AWARE OPTIMIZATION

A.I. Baranchikov, M.A.M. Abdi · 2025

The generation of high-quality synthetic data for relational databases has challenges regarding scalability, statistical accuracy, and the implementation of referential integrity. In this research study, we present a hybrid approach for synthetic data generation via an adaptive convex combination model to systematically integrate real and synthetic data. Relational schemas are characterized as CW complexes through sheaf theory to guarantee acyclic dependence resolution and to uphold foreign key constraints. This approach strengthens the enforcement of referential integrity with near-linear complexity and improves compute performance with GPU-accelerated θ-joins. The minimizing of Jensen-Shannon divergence ensures that synthetic data retains the structural and statistical properties of real data, hence constantly adjusting the blending factor. In comparison to existing methodologies, empirical validation with TPC-H and Postgres datasets demonstrates enhanced preservation of referential integrity, superior statistical integrity, and more computational scalability. This approach facilitates machine learning, database benchmarking, and privacy-preserving analytics.

Read the paper · More papers on PaperTik