Private Synthetic Data Generation for Mixed Type Datasets

Irene Tenison, A. I. Chen, Navpreet Singh, Omar Dahleh, Eliott Zemour, Lalana Kagal · 2024

In the face of escalating threats from privacy attacks on machine learning models, we propose a system that can artificially generate data that imitates real data but doesn’t contain any sensitive or personally identifiable information. The generated data, called synthetic data, will have the same semantic and statistical distribution as the original dataset but provide privacy guarantees. Compared to previous works that dealt with either structured or unstructured data separately, our work develops a complete hybrid pipeline for generating private synthetic datasets from complex datasets that consist of both structured (numerical or categorical) and unstructured data. The private synthetic data generated can be analyzed by collaborators and third parties without increasing the risks of leakage of sensitive data. We evaluate our system on Yelp reviews and drug side-effects datasets and calculate metrics for both quality and privacy. We introduce a context-aware exposure metric to quantify context-dependent memorization and use it along with exposure to evaluate privacy. Our evaluations demonstrate that our system generates meaningful private synthetic datasets that achieve good performance in characteristic similarity, utility, as well as privacy. Given these results, the generated synthetic data can be used by data scientists, researchers, and developers to address challenges related to data privacy, scarcity, diversity, and model testing in a wide range of applications including healthcare, insurance, and financial systems that rely on sensitive data.

Read the paper · More papers on PaperTik