Attribute Assailability and Sensitive Attribute Frequency based Data Generalization Algorithm for Privacy Preservation
Satish B Basapur, B S Shylaja, Venkatesh Venkatesh · 2021
The data publishing and analysis expose users’ and users group’s sensitive data and their identities. As a result, the adversary initiate inference attack using exposed users’ identifies and associated sensitive information. Some of the users’ attribute and users’ group attributes values are highly vulnerable and their contiguity increase privacy breach. The traditional k-anonymity based privacy models prevent identity disclosure but fails to prevent sensitive attribute values disclosure and ignore repercussion of assailable attributes on users’ group privacy. This research propose a novel data generalization algorithm that enhances users’ group privacy and data utility. The proposed method calculate the susceptibility of quasi-identifier attributes, frequency of sensitive attributes and their association. These values are used for form equivalence class that satisfies k-anonymity and β-likeness property. The proposed generalization algorithm considers both susceptibility, frequency of sensitive attributes and their association while anonymizing data. The extensive experiments are carried out on open source Apache Spark environment. The simulation results demonstrate that efficiency and effectiveness of proposed algorithm. This research work outperform other state-art data privacy model in terms of thwarting inference attack, information loss and accuracy of model.