HiDS Data clustering algorithm based on differential privacy
Shuhui Fang, Xuejun Wan, Jun Wang, Lin Chai, Wenlin Pan, Wu Wang · 2024
In recent years, the proliferation of high-dimensional sparse (HiDS) data has posed significant challenges to both data analysis and privacy protection. Traditional k-means algorithms often carry the risk of privacy leakage when applied to HiDS data, while differential privacy mechanisms offer effective protection for data privacy. Addressing such issues, this paper proposes a novel differential privacy k-means clustering algorithm. This algorithm first projects the data into a low-dimensional space, then privately generates a candidate center set, and finally performs privacy-preserving clustering on the candidate set. The algorithm proposed achieves the dual objectives of privacy protection and clustering analysis for HiDS data.