Probabilistic k-anonymity through microaggregation and data swapping
Jordi Soria-Comas, Josep Domingo‐Ferrer · 2012
k-Anonymity is a privacy property used to limit the risk of re-identification in a microdata set. A data set satisfying k-anonymity consists of groups of k records which are indistinguishable as far as their quasi-identifier attributes are concerned. Hence, the probability of re-identifying a record within a group is 1/k. We introduce the probabilistic k-anonymity property, which relaxes the indistinguishability requirement of k-anonymity and only requires that the probability of re-identification be the same as in k-anonymity. Two computational heuristics to achieve probabilistic k-anonymity based on data swapping are proposed: MDAV microaggregation on the quasi-identifiers plus swapping, and individual ranking microaggregation on individual confidential attributes plus swapping. We report experimental results, where we compare the utility of original, k-anonymous and probabilistically k-anonymous data.