A hybrid approach of privacy preserving data mining using suppression and perturbation techniques
Arshveer Kaur · 2017
In this epoch of growing technology the data collected by organizations has the requirement to preserve the privacy of the individuals. The techniques like anonymization, randomization are used to achieve the goal. But unfortunately anonymization leads to certain level of information loss while preserving privacy. Due to drift of data over the centralized server has added to the more demand of preserving privacy to dominate the trust of individuals. The application like the data of online shopping customers, patients, insurance is stored online over the centralized repository these days and it needs to maintain privacy of the individuals. In all the quoted applications the data contains many fields like email address, zip code, age, nationality etc. The quasi-identifiers like zip code, age, gender of a person does not seem to be very important to protect but these fields when linked with some other attributes can expose the identity or sensitive information of an individual. Thus the quasi-identifiers need a special scrutinity in the purpose of achieving privacy. The proposed hybrid approach combining suppression and perturbation for Privacy Preserving data mining takes care of these requirements. The method focuses on the goal of preserving privacy by suppressing and perturbing the quasi identifiers in the data of online shopping customers stored on centralized data repository without causing any loss to the information in the process. The method targets to overcome the limitation of information loss while preserving privacy. The experiment is carried by setting up a local server on the system and the simulation results are compared with anonymization to show that the goal to achieve privacy of quasi identifiers without information loss is successfully achieved.