Differentially Private Frequency Tables Based on Random Sampling

Takumi Sugiyama, Kazuhiro Minami · 2023

Previous research proposes a random sampling-based techniques [4], [16] to produce differentially private k anonymized data based on random sampling. However, their approach considers all variables in an original dataset as quasi-identifiers with full domain transformation to guarantee differential privacy. Since all the equivalence classes in anonymized data contain identical records with the same set of attribute values, that anonymized data is not useful for the analysis of microdata. However, we consider random sampling a promising approach to guaranteeing differential privacy on frequency tables since k-anonymized data in which all variables are treated as quasi-identifiers is equivalent to a frequency table.In this paper, we evaluate the feasibility of producing deferentially private frequency tables based on random sampling. Our experiments show that we can produce frequency tables of high data utility when we set small $\epsilon$ to choose a low sampling rate for a dataset with a large number of records.

Read the paper · More papers on PaperTik