Non-homogenous Slicing Anonymization with Subsequent Data utility Analysis for Privacy Preservation Data mining

Rajul Jain, Pranjali Singh · International Journal of Research in Advent Technology · 2019

With propulsion in the amount of data processed and released every day, privacy and security have become an indispensable factor in the data sphere.But data privacy and data utility seem to be in a constant tug-of-war with each other, with one factor having to compromise for the other.But if either utility or privacy is deprioritized beyond a certain point then the data might be rendered as either useless or vulnerable to severe privacy breaches.For this reason, it is essential to publish data in such a way that individual privacy remains intact, and the data is still useful for knowledge discovery, which is the main agenda behind Privacy-Preserving Data Mining (PPDM).This paper proposes a refinement of an existing PPDM technique known as slicing anonymization.Slicing has been previously proven to be an efficient technique for preserving the high quality of data while achieving high data privacy in publishing.In this paper, we target higher data utility and more secure data publishing using the concepts of probabilistic nonhomogenous suppression and attribute correlation.We also validate the results by comparing the pre-defined data quality metrics of the most used classification algorithms before and after applying this technique on the candidate dataset obtained from the Madhya Pradesh State Election Commission (MPSEC).The closeness of the results proves that our proposed algorithm maintains high data quality and ensures strong privacy preservation at the same time.

Read the paper · More papers on PaperTik