Research and Application Path Analysis of Deep Learning Differential Privacy Protection Method Based on Multiple Data Sources
Junhua Chen, Yiming Liu · Atlantis Highlights in Computer Sciences/Atlantis highlights in computer sciences · 2022
The deep learning model will contain user-sensitive information during training.When the model is applied, the attacker can recover the sensitive information in the training data set through model inversion attacks, and directly or indirectly disclose the user-sensitive information.The existing methods can not solve the problem that the privacy budget accumulates with the increase of training times.This paper proposes a method of research on deep learning differential privacy protection method based on multiple data sources, aim at making privacy consumption independent of the number of training epochs to guarantee the potential to work with large datasets.First, we calculate the privacy budget upper bound to optimal experiment selection for parameter estimation.Second, we use the upper bound to determine the number of group, also to balance the number of group and the data size of the subdataset, avoiding data relying on a single model causes leakage of user sensitive information.Finally, we ensemble several models with majority voting, and perturb single model the traditional convolutional deep belief network (CDBN) objective functions, to descend the dependence of privacy budgets on the training deep learning model and improve machine learning results.We applied our model to a health social network dataset and MNIST dataset, and the results show that our method has high privacy protection ability than the existing method for sensitive information on the training dataset.Moreover, standardization can be a feasible path for the generalized application of the technique, which is beneficial for the stability of the application of differential privacy protection techniques and the subsequent feedback updates.