Comparative Analysis of Utility Measures for Anonymized Data: Structured, Semi-Structured, and Unstructured Information in Microdata
Jundong Lee, Yongwan Joo, Jiyon Lee, Soon Seok Kim · Asia-pacific Journal of Convergent Research Interchange · 2024
This paper examines the utility measures of anonymized structured, semi-structured, and unstructured data, with a particular focus on the impact of anonymization on data quality.While anonymization is a key method for protecting personal information under the principles of Article 3 and Article 58-2 of the Personal Information Protection Act, it often creates a trade-off between data security and utility.The purpose of this study is to investigate the effect of anonymization on data utility, or information loss, in structured data (e.g., numerical and textual values), semi-structured transactional data, and unstructured image data.For structured data, this paper discusses traditional utility measures such as covariance, correlation, mean squared error, and absolute error, as well as utility measures within privacy protection models like k-anonymity.For semi-structured transactional data, it focuses on the k manonymity model and the key utility measure used in this model, the Normalized Certainty Penalty (NCP).For unstructured image data, the utility is evaluated using contemporary measures such as Fréchet Inception Distance (FID), Learned Perceptual Image Patch Similarity (LPIPS), and Structural Similarity Index Measure (SSIM) to assess the similarity between anonymized and original images.By thoroughly reviewing these measures, this paper aims to demonstrate that balancing anonymization and data utility remains a significant challenge for future research.Moreover, it seeks to contribute to organizations and businesses aiming to generate higher quality anonymized data for meaningful analysis.Future research will extend the discussion on utility measures to other forms of unstructured data, such as audio and text.