Examining the Utility of Differentially Private Synthetic Data Generated using Variational Autoencoder with TensorFlow Privacy
Bo-Chen Tai, Szu-Chuang Li, Yennun Huang, Pang-Chieh Wang · 2022
With the emergence of AI(artificial intelligence), it is becoming more and more critical for organizations to utilize it to their advantage. However, organizations that possess a decent amount of data might not have the technical competence to perform machine learning, and vice versa. Hence, it is reasonable for the two kinds of organizations to work together to realize the value of the data. With the increasing concern over data privacy, regulations such as GDPR(General Data Protection Regulation) prevent an organization from sharing data with another unless the data is processed to the point that the individuals in the data are not identifiable. Various ways of data anonymization have been proposed and developed, including the ones that utilize neural networks to achieve the goal, like AE, VAE, and GAN. With the addition of a differential privacy framework like TensorFlow Privacy, privacy can be guaranteed, but data still needs to be usable after privacy protection measures are deployed. The present study aims to integrate TensorFlow Privacy into the synthetic data generation process and evaluate its usefulness for daily use in the industries. Since TensorFlow Privacy brings a provable privacy guarantee to synthetic data, the present study focuses on the evaluation of data utility. TensorFlow is widely used for machine learning in the industry and academically. TensorFlow Privacy, which is also developed by Google, can prove to be a valuable addition to the synthetic data generation pipeline. The result shows that VAE with TensorFlow Privacy 1) generates synthetic data with good data utility in most cases in terms of descriptive statistics and machine learning classification tasks, and 2) The customizable TensorFlow Privacy parameters work as intended in terms of privacy-utility trade-off.