EFFICIENT LOCAL RECODING ANONYMIZATION FOR DATASETS WITHOUT ATTRIBUTE HIERARCHICAL STRUCTURE
Mohammad Rasool Sarrafi Aghdam, Noboru Sonehara · 2013
Privacy is one of the main concerns in data publishing especially when releasing datasets involving human subjects contain sensitive information. Hence, to protect the privacy of individuals, a model that is widely used for privacy preservation in publishing micro-data, is k-anonymity. It reduces the linking confidence between sensitive information and specific individual by 1/k ratio. However, k-anonymous dataset loses its accuracy due to the information loss. Most of the existing k-anonymization approaches suffer from huge information loss. In this paper we study the information loss issue and we propose a new model based on distance calculation between tuples including numerical and categorical attributes which is independent of attributes hierarchical structures. Then based on the proposed model we present the SpatialDistance (SD) heuristic algorithm for kanonymization. Our extensive study on real datasets shows that the proposed algorithm in comparison with existing well-known algorithms offers much higher data utility and reduces the information loss significantly. It also provides higher privacy protection for outliers.