Knowledge distillation-based privacy-preserving data analysis
Dong Huang, Xiucai Ye, Tetsuya Sakurai · 2022
Over the past decade, there has been an increasing focus on data privacy, which is one of the most popular areas of research today. Due to legal restrictions, it is difficult to centralize data, so multi-party collaborative learning becomes difficult to achieve. In this paper, we propose a data collaborative analysis method for protecting data privacy: Knowledge Distillation-based Analysis (KDDA). The proposed method can effectively solve the above problems. The framework of the proposed method is based on "teacher-student" network. Specifically, each institution (e.g., hospital) used sensitive data to train local models, which we call teacher models. Then we introduce unlabeled public datasets. We design an aggregator that aggregates all teacher models to generate pseudo-labels for the public dataset. It is worth noting that to minimize the risk of data privacy exposure, we limit the number of generated pseudo-labels. Finally, we train the student model using semi-supervised learning on the public dataset with pseudo-labels. Since the student model is not directly trained on sensitive data, it will not lead to the leakage of data privacy, so it can be publicized. To evaluate the effectiveness of the proposed method, we conduct experiments on three popular image benchmark datasets: MNIST, SVHN, and CIFAR-10. Experimental results show that the proposed method is effective.