Phantom Anonymization: Adversarial testing for membership inference risks in anonymized health data

Thierry Meurers, Mehmed Halilovic, Karen Otte, Jérémie Despraz, Bayrem Kaabachi, Bogdan Kulynych, Jean Louis Raisaro, Fabian Praßer · Computers in Biology and Medicine · 2025

OBJECTIVE: In medical research, where datasets often describe individuals with specific medical conditions, revealing someone's membership in a dataset can lead to significant privacy breaches. Anonymization is an important technique to protect health data by modifying it to reduce privacy risks. However, quantifying membership inference risk is challenging, which we address with our work. METHODS: We propose a framework to quantify residual membership inference risks in structured, tabular data. It adapts techniques previously utilized in the assessment of synthetization methods, where artificial data is generated by a model reproducing the statistical properties of the original data. The core idea involves training a classifier designed to detect the presence of target records in anonymized datasets. The classifier is trained with data samples that have been anonymized in the same manner as the dataset under attack. To evaluate the framework, we conducted experiments across different anonymization approaches and adversarial scenarios. RESULTS: Our results indicate that the approach is suitable for detecting residual privacy risks in anonymized data. A comparison across different anonymization methods shows that their effectiveness depends not only on the chosen privacy model, i.e. on the approach with which privacy risks are measured, but also significantly on how data is modified to meet predefined risk thresholds. CONCLUSION: Our framework provides an empirical way to assess residual membership inference risks for a wide range of anonymization methods. As it adopts a technique developed for synthetic data, it also enables comparisons of residual risks between synthetic and anonymized datasets.

Read the paper · More papers on PaperTik