Clustering Dissimilar Tuples: A Stronger Notion of Privacy
International Journal of Emerging Trends in Engineering Research · 2020
Identity disclosure and attribute disclosure have always been a major concern while publishing data.k-anonymity tries to solve identity disclosure but doesn't prevent attribute disclosure which leads to homogeneity and background knowledge attack.Preserving privacy of an individual is becoming more challenging due to increasing number of homogeneity and background knowledge attacks.l-diversity model has been proposed to thwart these attacks but it doesn't fulfil its obligations.Several authors found l-diversity model to be inadequate, hence they put forth another model called t-closeness.Over the years, many investigations and experimentations conducted by various researchers shows that t-closeness does not provide a clear relationship between the threshold value t and information gain and it also shows that Earth mover's distance, a distance metric used by t-closeness model, becomes complex with multiple sensitive attributes.In view of this challenge, we propose a stronger notion of privacy called Clustering Dissimilar Tuples (CDT) to thwart homogeneity and background knowledge attack by formalizing the idea of processing the original dataset initially wherever these attacks possibly occur.These attacks are found to occur in the tuples of sensitive attributes.Hence CDT processes the tuples of sensitive attributes to form equivalence classes consisting of dissimilar tuples.Through experimental evaluations, we show that CDT is practical and can be implemented efficiently with minimum utility loss and maximum privacy gain.