Two-phase entropy based approach to big data anonymization

Ashish Ranjan, Prabhat Ranjan · 2016

Data outbreak following the flourishing of technologies like cloud computing, social networks, and many more stimulate emergence of new paradigm called “Big Data”, which not only improved the decision-making, but also provided a means to easy privacy violation of an individual. Existing data anonymization techniques adhering to certain privacy model such as k-anonymity or l-diversity proved their worth, but demands identifying potential quasi-identifier set manually by the domain expert which is very much prone to human error. The paper focus on the problems of choosing appropriate quasi-identifier set and minimizing information loss due to anonymization process. The proposed approach is two phase entropy driven approach to find the quasi-identifier set in a first phase, which is extended to second phase to anonymize the dataset based on selected quasi-identifier set. Our approach ensures that dataset maintains its statistical properties along with minimal information loss and better privacy.

Read the paper · More papers on PaperTik