Anonymization of Bigdata using ARX Tools

RK Shyamasundar, Manoj Kumar Maurya · 2024

Anonymization and sanitization of data has become extremely important in the context of various privacy laws around the world. K-Anonymity is a widely used technique that is a property for the measurement, management, and governance of the data anonymization. There is a loss of information in most implementations of K-anonymity and are not practically usable over large datasets with a number of attributes. In this paper, we explore a practical way of anonymization and sanitization using open-source solutions such as ARX tools. We explore K-anonymization features like optimal, diversity, and closeness, and models of privacy, to realize anonymization with optimal or minimal loss of information. We use public US census data for our study and use metrics, like information loss, utility, and privacy. As there is a need to strike a balance between minimizing information loss and maximizing utility in realizing privacy, the study employs an ensemble of algorithms of K-anonymity with/without ℓ-diversity, and t-closeness. The experimental results demonstrate that the combined approach outperforms ϵ-differential privacy among different algorithmic combinations on different parameters. Further, an innovative approach for multi-quasi identifier datasets (DS) is proposed to enhance the utility of K-anonymization by integrating it with differential privacy(DP) either local or global; it makes the database (DB) rows quite independent that is one of the main prerequisites for applying DP.

Read the paper · More papers on PaperTik