Distributed Data Anonymization
Fabio Martinelli, Mina Sheikhalishahi · 2019
Data generalization is a widely-used privacy technique, in which an accurate value of sensitive information is replaced with a more general representation. The problem of data generalization becomes challenging when data is distributed among several agents, who are interested in releasing their table of data to shape a data mining algorithm on the whole of their data. The main issue originates from the fact that when each agent generalizes her own dataset locally, the released tables of data suffer from non-homogeneity. To sole the issue, all agents can generalize their data to the widest range of generalization. However, this approach causes utility loss. To optimally address this problem, in this study we present a framework that serves as a tool for data owners to generalize their data homogeneously before being published. The effectiveness of the proposed mechanism is validated through an experimental analysis on a benchmark dataset.