Privacy preservation in k-means clustering by cluster rotation
S. S. Shivaji Dhiraj, Ameer M. Asif Khan, Wajhiulla Khan, Ajay Challagalla · 2009
The use of clustering as a data analysis tool has raised concerns about the violation of individual privacy. This paper proposes a data perturbation technique for privacy preservation in k-means clustering. Data objects that have been partitioned into clusters using k-means clustering are perturbed by performing geometric transformations on the clusters in such a way that the object membership of each cluster and orientation of objects within a cluster remain the same. This geometric transformation is achieved through cluster rotation, i.e., every cluster is rotated about its own centroid. The clusters are first displaced away from the mean of the entire dataset so that no two clusters overlap after the subsequent cluster rotation. We analyze the privacy measure offered by this data perturbation technique and prove that a dataset perturbed by this method cannot be easily reverse engineered, yet is still relevant for cluster analysis.