A scalable approach for anonymization using top down specialization and randomization for security

S Athiramol, S. Sarju · 2017

Anonymization is a process of hiding the information such that an illegitimate user could not infer anything from the records, on the other hand an analyzer will get necessary information. That is the metrics that determines the goodness of an anonymization algorithm are data utility or information loss and data security. Although there are different algorithms that exist for the purpose of anonymization, all of them are having disadvantages in terms of utility, security, and also the execution time it requires. Here an approach is proposed that works for both numerical and categorical data in a time efficient manner. The system is built on top of Hadoop MapReduce framework. Although for any algorithm, 100 percent data utility and 100 percent security is not a promise, we could propose an algorithm that optimizes the result. Unlike other algorithms that suppresses all the records that cannot be anonymized, this paper suggests an algorithm called randomization that can be applied on it. This approach reduces the chance for background knowledge attacks that many algorithms failed to avoid.

Read the paper · More papers on PaperTik