Differentially Private Synthetic Data Generation Using Context-Aware GANs
Anantaa Kotal, Anupam Joshi · 2024
The widespread use of big data across various sectors has brought significant privacy concerns, particularly when sensitive information is shared or analyzed. Regulations like GDPR and HIPAA impose strict controls on handling data, making it difficult to balance the need for insights with privacy requirements. Synthetic data offers a promising solution, enabling the creation of artificial datasets that mirror real-world patterns without exposing sensitive information. For instance, synthetic data can simulate patient records or network flows for training machine learning models to conduct research without violating privacy laws. However, traditional synthetic data generation methods often fail to capture complex, implicit rules that relate different elements of the data and are essential in specific domains like healthcare. While these methods might replicate explicit patterns from the training data, they often overlook domain-specific rules that are not directly stated but are critical for maintaining realism and utility. For example, prescription guidelines, such as avoiding certain medications for patients with specific conditions or preventing harmful drug interactions, may not be explicitly represented in the original data. Synthetic data generated without accounting for these implicit rules can lead to medically inappropriate or unrealistic patient profiles. To address these limitations, we propose a framework called Context-Aware Differentially Private Generative Adversarial Network (ContextGAN). Our framework integrates domain-specific rules using a constraint matrix that explicitly encodes both explicit and implicit domain knowledge. The constraint-aware discriminator evaluates synthetic data against these rules, ensuring the generated data adheres to domain constraints. Furthermore, the discriminator is differentially private, ensuring privacy preservation by protecting sensitive details from the original data. We validate ContextGAN across multiple domains, including healthcare, security, and finance, demonstrating that it produces high-quality synthetic data that respects domain-specific rules while preserving privacy. Our results show that ContextGAN significantly improves the realism and utility of synthetic data by enforcing domain constraints, making it suitable for use in scenarios requiring both compliance with explicit patterns and implicit rules, all under strict privacy guarantees.