Privacy Preservation Measure using t-closeness with combined l-diversity and k-anonymity
Anil Prakash Dangi, Ravindar Mogili · 2012
Abstract- Public survey data that may Increase the exposure of privacy and census information about the particulars is called the Data sensitivity. To maintain the privacy increase the similarity in the data item and introduce redundancy in such a way that information about individual users can not be disclose. This technique also desires the actual information of the data to not change. In this work we propose a unique method by combining two of the most widely used privacy preservation techniques: K-anonymity and l-diversity. The k-anonymity privacy requirement for publishing micro data requires that each equivalence class (i.e., a set of records that are indistinguishable from each other with respect to certain “identifying ” attributes) contains at least k records. Diversity requires that each equivalence class has at least well-represented (in Section 2) values for each sensitive attribute. In this article, we show that-diversity has a number of limitations. In particular, it is neither necessary nor sufficient to prevent attribute disclosure. Motivated by these limitations, we propose a new notion of privacy called “closeness”. We first present the base model t-closeness, which requires that the distribution of a sensitive attribute in any equivalence class is close to the distribution of the attribute in the overall table (i.e., the distance between the two distributions should be no more than a threshold t). Based on entropy based closeness and distance measure between the class of data we propose a comprehensive technique to change the dataset to preserve the privacy while keeping the original meaning intact.