A Novel Mixed Integer Programming Formulation for Data Perturbation
Ting Zhang, Xinglu Liu, Jiguang Wang, Wai Kin Victor Chan · 2022
Data leakage will lead to privacy exposure and cause the loss of data assets and brand reputation. Data perturbation is an efficient approach to overcome the aforementioned challenge, which can protect the privacy by introducing noise into the original data or transforming the raw data. Existing literature perform data perturbation with random perturbation, clustering, generalization, suppression, anatomization, permutation, etc. The results offered by these methods are feasible solutions, whereas the optimality guarantee is not provided. In this work, we propose a novel mixed integer programming (MIP) formulation for the data perturbation problem that aims at minimizing the change of the mean and the variance. The proposed MIP model is able to yield optimal solution with respect to multiple constraints. Since the objective of the proposed formulation incorporates the distribution relevant information, it is of high flexibility and can control the range of the noise. Results reveal that our model outperforms the existing approaches in most of the tested datasets.