Greedy Clustering with Sample-Based Heuristics for K-Anonymisation
Grigorios Loukides, Jianhua Shao · 2007
Developing techniques for k-anonymising data has received much recent attention from the database research community. Good k-anonymisations should retain data utility and preserve privacy, but these are conflicting requirements and can only be traded-off. A method proposed recently attempted to achieve a balance between these two requirements, but its efficiency and effectiveness depend heavily on several empirically set parameters. In this paper, we propose sampling-based heuristics to optimally set up these parameters. We test the effectiveness of our methods by evaluating anonymisations in terms of accuracy in query answering and ability to prevent linking attacks.