A Non-Parametric Model for Accurate and Provably Private Synthetic Data Sets

Jordi Soria-Comas, Josep Domingo‐Ferrer · 2017

Generating synthetic data is a well-known option to limit disclosure risk in sensitive data releases. The usual approach is to build a model for the population and then generate a synthetic data set solely based on the model. We argue that building an accurate population model is difficult and we propose instead to approximate the original data as closely as privacy constraints permit. To enforce an ex ante privacy level when generating synthetic data, we introduce a new privacy model called ϵ synthetic privacy. Then, we describe a synthetic data generation method that satisfies ϵ-synthetic privacy. Finally, we evaluate the utility of the synthetic data generated with our method.

Read the paper · More papers on PaperTik