Research on Data Security Protection Method Based on Improved K-means Clustering Algorithm
Jinghui Cai, Dashun Liao, Jianqiu Chen, Chen Xue, Tingting Liu, Jialin Xi · 2020
Data security is a severe challenges in the era of Internet. The traditional K-means clustering algorithm can protect the data security properly when it is applied in the field of information security, but the accuracy and availability of its clustering process are difficult to meet the needs of reality. In order to effectively achieve the high availability and accuracy of clustering algorithm, aiming at the shortcomings of traditional K-means algorithm, this paper studies the improvement of K-means clustering algorithm based on differential privacy protection. The core theory of K-modes is used to select the initial point of clustering, and the shortest distance from the current point to the original cluster center point is found by using Euclidean distance, so as to get the cluster again. The differential privacy protection technology is introduced and improved based on Laplace mechanism. By calculating the distance between the data sample and the center point, the specific location of sensitive attributes in the data sample is obtained, so as to change the order of adding noise and obtain more accurate clustering results. The experimental results show that compared with other data security protection algorithms, the algorithm proposed in this paper has obvious advantages in clustering effect, clustering accuracy and time complexity.