A Study on The Improvement of Information Loss Metrics in Real-Time Stream Data Anonymization
Ji Yeon Lee, Yong Wan Joo, He Young Yun, Soon Seok Kim · Asia-pacific Journal of Convergent Research Interchange · 2023
Stream data refers to information collected in real time, such as crime report information, online sales transaction information, and information from patient monitoring devices in hospitals.This paper deals with the anonymization problem, which is the privacy issue of stream data.Usually, the main consideration in anonymization is how to process the data in such a way that it is secure and useful at the same time.Data usefulness refers to the quality of data and is measured by a so-called loss of information measure.In this paper, we review the current information loss metric in real-time stream data anonymization and propose a new metric that improves the disadvantages by applying Goldberger et al.'s scheme.The measure proposed by Goldberger et al. is characterized by considering the total equivalent class of the data table and the number of records in the equivalent class, which were not previously applied in the generalized dataset to which the k-anonymity model was applied.In the past, only one equivalent class within one cluster was assumed, so it was not included in the consideration.The reason is that records that do not satisfy k-anonymity within one cluster are allocated to other clusters that can be moved, that is, clustered with the next-rank quasi-identifier, or are deleted if there are no clusters that can be allocated.However, in terms of information loss, it is reasonable to assign them to other clusters because deleting those records leads to increased loss.Therefore, it is desirable to minimize information loss by allocating these records to another cluster that can be moved or, if there is no cluster to be assigned, allocating all of these records to a separate independent cluster instead of deleting them.The proposed idea focuses on the latter rather than the former.In this case, it is because several equivalent classes with quasi-identifiers can exist in an independent cluster.Therefore, by using the proposed metric of Goldberger et al., it is necessary to generalize to minimize the measured loss value after measuring the information loss by considering the equivalent class and the number of records in the equivalent class.