A new technique ensuring privacy in big data: Variable t-closeness for sensitive numerical attributes
Zakariae El Ouazzani, Hanan El Bakkali · 2017
Due to the massive growth of data, the notion of big data has obviously gained momentum in recent years. Thus, an enormous amount of personal information could be contained in high dimensional data sets which requires the preservation of privacy before publishing such information. In this context, several anonymization techniques, such as generalization and randomization, have been designed in order to sanitize the data and consequently ensure privacy in big data. Over the years, t-closeness has been treated with great interest as an anonymization technique ensuring privacy in big data. Despite the fact that many algorithms for t-closeness have been proposed, many of them admit that the threshold t of t-closeness is set to a fixed value. Here, a novel way in applying t-closeness for a sensitive numerical attribute is presented. A new algorithm called variable t-closeness for sensitive numerical attributes was proposed. Our proposed algorithm gives good results in terms of data anonymization. This algorithm was experimentally evaluated on a test table. Furthermore, we highlighted all the steps of our proposed algorithm with detailed comments.