Preserving Privacy in Fine-Grained Data Distillation With Sparse Answers for Efficient Edge Computing
Ke Pan, Maoguo Gong, Kaiyuan Feng, Hui Li · IEEE Internet of Things Journal · 2024
In the field of Internet of Things (IoT), data distillation has been thought of as a key method to condense the original real dataset into a tiny synthetic dataset with less training burden while maintaining as much data utility as possible for training deep learning models. However, the data synthesis process may remember some sensitive information about the original dataset, which may raise privacy concerns for data owners. To address this problem, we present a novel differential privacy (DP)-based data distillation algorithm. Specifically, in the data distillation phase, we first randomly pick a training model from the model pool in each epoch, and then build a fine-grained distribution matching to generate informative data for improving the task-oriented model performance. In the privacy preservation phase, we selectively perturb input features that are more important for model training based on the sparse vector technique to protect the sensitive information contained in the original dataset and reduce privacy costs. Extensive experiments across several real-world datasets demonstrate that our algorithm can achieve higher data utility and model accuracy than existing solutions.