Mixed data k -Anonymization by Consistent Maximal Association and Microaggregation
Julien Ah-Pine, Nathaniel Gbenro · 2025
This paper addresses the challenge of anonymizing mixed data, comprising both categorical (qualitative) and numerical (continuous) variables, while preserving data utility. The inherent heterogeneity of such data complicates the use of traditional anonymization methods. To overcome this limitation, we propose a novel microaggregation-based framework for k-anonymization that integrates statistical association measures applicable to both variable types, ensuring a coherent and consistent treatment. Our approach, called Mix-R 2, relies on a unified set of core concepts grounded in analysis of variance, enabling the application of a common methodology to both categorical and numerical attributes. By leveraging these consistent association measures, the framework improves the robustness of the k-anonymization process, delivering strong privacy protection while maintaining high data utility. Numerical experiments on benchmark datasets demonstrate the effectiveness and advantages of our method, highlighting its contribution to privacy-preserving analysis of mixed-type data.