Personalized Privacy-Preserving Data Utilization Approach Powered by Distributed-GAN
Shuo Wang, Chao Wang, Tianshuo Dong, Yunhua He, Ke Xiao · Big Data Mining and Analytics · 2024
Recently, Machine Learning (ML) has made great achievements in a wide range of fields with the increasing data exploring. However, the development of ML faces the challenges of data silos due to privacy considerations. Then with distributed ML architecture appearing, the training tasks can be completed without sharing data. In particular, distributed Generative Adversarial Network (GAN), as a generative model, can rebuild dataset for downstream tasks and does not need to re-access users' private datasets even if the target task changes. However, the training of distributed-GAN still faces serious privacy threats. Existing work usually adopts a uniform level of privacy protection, which does not meet the requirements about personalized privacy protection for each user. In this paper, we propose a privacy-preserving training framework of distributed-GAN, which combines differential privacy to provide users with personalized privacy protection. Further, a privacy-driven incentive strategy is proposed, which designs a series of smart contracts to provide customized payments for data owners with different privacy preferences as compensation for their privacy cost. At last, we conduct sufficient privacy analysis and experimental validation, which demonstrate that our approach can optimize generative model with lower privacy cost and generate higher quality data for downstream tasks.