Federated privacy preservation algorithm combining conditional generative adversarial network and knowledge distillation
Zilong Wang, Shi Lin, Chuan Hong Zhou · Egyptian Informatics Journal · 2025
In the digital era, the value of personal data has become increasingly prominent, but with it, the risk of privacy leakage has intensified. At present, there is a balance between the robustness of privacy protection mechanisms and the efficient utilization of data, as well as the need for a comprehensive performance evaluation system. Federated learning has become a critical technology for solving the contradiction between privacy protection and model optimization because it allows multiple parties to jointly train models without sharing original data. However, traditional federated learning still faces many challenges, especially when dealing with susceptible information or small sample data sets, and the balance between model generalization ability and privacy protection strength is difficult to grasp. To this end, this study proposes a federated privacy protection algorithm that combines conditional generative adversarial network and knowledge distillation, aiming to improve model training efficiency while minimizing the risk of data leakage. The introduction of CGANs gives the model more robust generative capabilities, enabling the creation of high-fidelity synthetic data for local training, thereby reducing the reliance on accurate data. The experimental results show that the accuracy rate of the model trained with synthetic data reaches 89.2%, which is only 0.8% lower than that of using accurate data but significantly reduces the possibility of data leakage. At the same time, the application of KD technology promotes knowledge transfer between models. Allowing the student model to imitate the behaviour of the teacher model can still maintain high learning efficiency even under limited data conditions. In a cross-device federated learning experiment, the student model optimized by KD performed comparably on the test set as a large model trained on a single device, with accuracy rates of 91.4% and 91.7%, respectively, while the former required only about half the time and computing resources of the latter.