Application of privacy-preserving clustering methods using homomorphic encryption algorithms

Ismail Aydin · Istanbul Şehir University Institutional Repository (Istanbul Şehir University) · 2018

The need of protection and processing of the sensitive data in large scale data systems (for example data derived from nancial systems, militaristic systems or social media platforms) is a common problem. Usage of traditional cryptographic methods for data protection mainly needs at least two of the ciphering, deciphering and data processing works to be done on the same side. Because of this, with increase of the data size there will be a need for higher processing power to work on the data. Using traditional encryption algorithms for protection of the sensitive data on large scale systems, also brings the need of exchanging the needed keys for protection and processing the data. Homomorphic encryption schemes have enough exibility that, they should be used on data systems that contains data from multiple parts, because of its feature of allowing to process the encrypted data like its non-encrypted form. With the usage of homomorphic encryption schemes and proper data learning systems on encrypted data, distribution of sensitive data to dierent parties can be done without violatingitsprivacy. Inthisthesis, weproposeamethodtorunmathematicalcomputations which needs high processing power on a common platform which oers high processing power of data but not on parties that the sensitive data will be distributed. As a result the partners of this systems will not need to have high processing power to function on the data because the high processing demanding tasks would be done on the common platform. In this research Paillier Cryptographic system was used to protect data privacy. Paillier Cryptographic algorithm's most prominent features are its asymmetrical and partially homomorphic behavior. We proposed a system that uses privacy preserving distance matrix calculation as input for several clustering algorithms which are commonly used in machine learning systems. Our system is evaluated considering dierent data lengths and dierent key lengths. Four dierent data clustering methods have been tested. By applying clustering algorithms on both encrypted and plain forms of the same data for dierent key and data lengths, we obtained performance results by using six dierent metrics.

Read the paper · More papers on PaperTik