A Study on Improving Metrics of Information Loss in Anonymization of Semi-structured Transaction Data with Set-valued Attributes
Ji Yeon Lee, Dong Hyun Kim, Young Ae Lee, Soon Seok Kim · Asia-pacific Journal of Convergent Research Interchange · 2024
Structured data typically refers to data composed of a single character or number within a single cell of a table.If, instead, a single cell contains a set of multiple characters or numbers, we refer to this as semi-structured data.For example, the list of items purchased at a supermarket or diagnoses for a patient in a hospital represents such cases.In these cases, the list of values for an individual is called a transaction.We address the anonymization problem for privacy protection in semi-structured transaction datasets composed of these values.Specifically, lists of purchased items or diagnoses contain sensitive information that must be protected for individual privacy.Regarding this anonymization problem, we previously proposed an improvement to the existing Local Generalization (LG) algorithm, resulting in the new Local Generalization & Reallocation (LGR) algorithm.However, the data must be secure and valuable from the perspective of utilizing or analyzing anonymized personal data.This implies that data quality must be preserved to facilitate analysis, which means that information loss during anonymization must be minimized.Our existing LGR algorithm used the Normalized Certainty Penalty (NCP) as a measure for calculating information loss, the same as the existing LG algorithm.The NCP measure calculates the ratio of generalized items to the total number of items.While this measure is simple to compute and easily applicable to various datasets, it has the drawback of potentially high information loss, which can reduce data utility.To address this drawback, we propose a new Information Gain-based Heuristic (IGH) measure and aim to verify its effectiveness theoretically.The proposed measure has the advantage of minimizing information loss and maximizing data utility compared to the existing NCP method.