Proposal of a New Anonymization Algorithm for Real-Time Stream Data
Sung Hyun Hong, Gyu Sung Lee, D. Kim, Soon Seok Kim · Asia-pacific Journal of Convergent Research Interchange · 2023
Real-time stream data refers to data which collected in real time such as personal vital signs information collected from various PoCs (Point of Care, point-of-care medical equipment) in hospitals, real-time crime report information, and online sales transaction.We propose a new anonymization algorithm for privacy protection in these real-time collected stream data.The proposed algorithm significantly improved performance in four aspects compared to the UBDSA algorithm proposed by Ugur and Osman.First, after forming a generalization tree, precompute was performed to measure information loss in advance.Second, whenever each transaction data (record) is entered, quasi-identifier columns and non-identifier columns are separated and stored through slicing, and serial numbers are attached to each slice separated at the time of separation for combine before publishing, That is, when clustering for each transaction, only quasi-identifier columns are stored and generalized to reduce storage space and improve calculation performance.In particular, it is more efficient for large numbers of columns or large-capacity data.Third, the performance in the cluster assignment (AssignCluster) step was improved as follows.First, from the initial cluster assignment process below the delay threshold (the maximum allowable number of transaction records from input to publishing), clustering was performed in consideration of information loss, rather than assigning to separate clusters, unlike existing algorithms.Finally, performance was improved in generalization and publishing (Publish stage).First, after forming a generalization tree, precompute was performed to measure information loss in advance.Second, whenever each transaction data (record) is entered, quasi-identifier columns and non-identifier columns are separated and stored through slicing, and serial numbers are attached to each slice separated at the time of separation for combine before publishing, That is, when clustering for each transaction, only quasi-identifier columns are stored and generalized to reduce storage space and improve calculation performance.In particular, it is more efficient for large numbers of columns or large-capacity data.Third, the performance in the cluster assignment (AssignCluster) step was improved as follows.First, from the initial cluster assignment process below the delay threshold (the maximum allowable number of transaction records from input to publishing), clustering was performed in consideration of information loss, rather than assigning to