GetFed: Accurate, Differentially Private Federated Learning With GAN-Based Data Generation

Hao Bai, Kun He, Yuqing Li, Jing Chen, Haowei Li, Zhongmou Liu, Xuanang Yang, Ruiying Du · IEEE Transactions on Dependable and Secure Computing · 2025

Federated Learning (FL) aims to train neural network models using distributed data resources from multiple clients without sharing raw data. One of the key challenges in FL is non-independent and identically distributed (non-IID) data, which may affect model accuracy. To address this issue, some schemes leverage Generative Adversarial Networks (GANs) to generate virtual data and combine it with the real data to achieve a balanced data distribution. However, there are risks of privacy leakage from the collected virtual data and aggregated gradients. In this paper, we propose GetFed, an accurate and differentially private FL framework with GAN-based Data Generation on non-IID Data. We integrate Differential Privacy (DP) into the GAN training and federated aggregation phases to prevent clients’ privacy leakage. To balance privacy and accuracy, we first design a privacy-preserving virtual sample generation algorithm for GAN training that dynamically reduces unnecessary noise as the quality of virtual samples improves. Additionally, we design an adaptive DP-based secure aggregation algorithm that decreases the added noise as the model approaches convergence. Furthermore, we implement a real-virtual ensemble training algorithm, employing an ensemble learning strategy to better mix virtual and real samples for enhanced global model accuracy. This approach ensures clients benefit from both the authenticity of real samples and the balanced data distribution provided by virtual samples, effectively mitigating the data heterogeneity inherent in non-IID scenarios. Extensive experiments demonstrate that compared with state-of-the-art schemes, GetFedimproves model accuracy by 6–47% and reduces training time by 50%.

Read the paper · More papers on PaperTik