Hybrid SVD BASED DATA TRANSFORMATION METHODS FOR PRIVACY PRESERVING
Syed Syed, Tarique Ahmad, Shameemul Haque, Shoeb Ahmed Khan · 2014
At the present time privacy issues are main concern for many government and other private organizations to delve important information from large repositories of data. Privacy preserving clustering which is one of the techniques emerged to addresses the problem of extracting useful clustering patterns from distorted data without accessing the original data directly. In this paper a hybrid data transformation method is proposed for privacy preserving clustering in centralized database environment. The proposed hybrid method takes the advantage of two existing techniques such as Singular Value Decomposition (SVD) and shearing based data perturbation. Experimental results demonstrate that the proposed method efficiently protects the private data of individuals and retains the important information for clustering analysis. ATA collection has increased rapidly with the evolution of internet and information technology. In order to en- hance their business, organizations are using data mining for extracting new patterns and relationships. The process of sifting through huge databases, for extracting useful, hidden patterns is called as data mining. The techniques of data min- ing are widely used in business and scientific communities such as medical, healthcare, insurance, banking, marketing etc. Association rules, classification, clustering, regression are some of the data mining tasks. Cluster analysis partitions data into several categories or useful groups (clusters) based on the similarity in the data. It is an unsupervised learning method, which is used for the exploration of inter relationships among a collection of patterns, by organizing them into homogeneous clusters. Relative distance or relative density between the ob- jects is taken as the similarity measure for the clustering ob- jects. Clustering is performed based on the principle of max- imizing the intracluster similarity and minimizing the inter- cluster similarity. To resolve the problem of privacy, a new research area called privacy preserving data mining has been evolved. The process of privacy preserving data mining is to extract useful patterns without breaching the privacy of individuals. Differ- ent techniques have been proposed for protecting the privacy of individuals such as data modification, data partitioning, data restriction and data ownership (1). Data mining is providing numerous benefits, there is a negative impact with data mining is the risk of privacy invasion. This problem is addressed by a new branch of data mining which takes the privacy issues under consideration is known as privacy pre- serving data mining. The goals of privacy preserving data mining are A. Protection of privacy in data release B. Privacy is protected among multiple collaborating parties C. Protecting the sensitive knowledge patterns ex- tracted with data mining tools. Privacy preserving clustering methods can extract valid clustering patterns without breaching the privacy of individu- als. Different approaches have been developed to effectively shield the sensitive information contained in databases such as access control, perturbation techniques, anonymity, and se- cure multi-party computation. In this paper a hybrid data transformation method is proposed for privacy preserving clustering, which is a combination of Singular Value Decom- position (SVD) and shearing based data perturbation.