A Unified Clustering-Based Anonymization for Privacy-Preserving Data Publishing with Multidimensional Privacy Quantification
Anselme Herman Eyeleko, Tao Feng, Yan Yan · Information · 2026
As widely adopted privacy models in privacy-preserving data publishing (PPDP), k-anonymity and ℓ-diversity have been extensively studied by researchers to enable the release of useful information while preserving data privacy. However, existing methods suffer from several limitations. They often rely on single-dimensional privacy models and lack unified metrics for accurately quantifying privacy leakages. Many approaches overlook the impact of semantic similarity and adversarial prior and posterior beliefs among sensitive attributes and frequently employ suboptimal similarity measures that fail to account for the heterogeneous nature of quasi-identifiers, thereby degrading both privacy protection and data utility. To address these challenges, this paper proposes CAMDP, a unified clustering-based anonymization method for privacy-preserving data publishing with multidimensional privacy quantification. CAMDP constructs equivalence classes that satisfy k-anonymity while simultaneously enhancing sensitive attribute diversity, reducing semantic similarity, and limiting divergence between prior and posterior adversarial beliefs. A unified multidimensional metric is introduced to jointly quantify privacy leakage and information loss, guiding the anonymization process. Additionally, a similarity-aware distance metric tailored to mixed-type quasi-identifiers is employed to reduce information loss. Experimental results on three benchmark datasets, Adult, Careplans, and Airline, demonstrate that CAMDP consistently outperforms state-of-the-art methods. Across all tested configurations, CAMDP achieves the lowest average privacy leakage (0.1235, 0.0795, and 0.1855, respectively), lower average information loss (0.626, 0.636, and 0.60, respectively), and the lowest average intra-cluster dissimilarity (0.586, 0.635, and 0.573, respectively), while maintaining competitive execution time across the three datasets.