Quantifying the Effects of Anonymization Techniques Over Micro-Databases
Debanjan Sadhya, Bodhi Chakraborty · IEEE Transactions on Emerging Topics in Computing · 2022
Micro-databases are unique datasets that contain person-specific information about individuals. Preserving the privacy of such datasets has become a cause for serious concern since this massive repository of personalized data regularly gets published in the public domain. Sanitization mechanisms are specialized techniques that provide the required privacy guarantees to the published data. The work in this article establishes an efficient framework for quantitatively estimating the effectiveness of any privacy-preservation scheme which employs the anonymization principle. In our study, we have introduced an information-theoretic metric termed asSanitization Degree$(\eta)$which assigns a cumulative score in the range [0,1] for a generic anonymization process. The design of our proposed metric is based on the fundamental fact that any sanitization mechanism attempts to reduce the amount of correlated information within the database attributes while simultaneously preserving the utility of the original dataset. Furthermore, we have characterized theprivacy-utilitytradeoff associated with our model by establishing a working relationship between these two fundamental quantities. We have empirically computed the value of$\eta$for three popular anonymization models ($k$-anonymity,$l$-diversity, and$t$-closeness) over a couple of real-life micro-databases, thereby demonstrating the practicability of our study.