Privacy Playground: A Framework for Assessment of Data Anonymization Methods
B John, Mahmoud Mahmoud, Abdolhossein Sarrafzadeh, Ahmad Patooghy · 2025
Configuring privacy-preserving methods for real-world datasets remains a significant challenge, as optimal parameters vary substantially across different data characteristics and use cases, requiring careful balancing between privacy protection and data utility. Current approaches rely heavily on manual parameter tuning, leading to inefficient and potentially suboptimal privacy-utility trade-offs. To address this critical gap, we present a novel framework that automatically assesses privacypreserving techniques across diverse datasets to set them at the right working spots. Our framework employs statistical analysis to narrow the parameter search space, dramatically reducing the computational overhead typically associated with finding optimal privacy configurations while maintaining high utility. The framework enables data scientists to rapidly identify optimal privacy-utility trade-offs for any privacy-preserving method. We demonstrate the framework's effectiveness by evaluating Differential Privacy (DP) and K-Anonymity (KA) methods on a wide range of input data, i.e., bank, disease, and oil well datasets. Through our experimentation, we observed that DP achieves optimal performance at moderate privacy budgets$(\epsilon$values between 0.01 and 0.0358), while KA's effectiveness varies with dataset characteristics, favoring lower$k$values for smaller datasets and higher$k$values for larger ones.