Anonymity in data publishing and distribution
David J. DeWitt, Kristen LeFevre · 2007
Numerous organizations collect and distribute non-aggregate personal data for a variety of different purposes, including demographic and public health research. In these situations, the data distributor is often faced with a quandary: On one hand, it is important to protect the anonymity and personal information of individuals. One the other hand, it is also important to preserve the utility of the data for research. This thesis presents an extensive study of this problem. We focus primarily on notions of anonymity that are defined with respect to individual identity, or with respect to the value of a sensitive attribute. We propose a variety of techniques that use generalization (also called recoding) to produce a sanitized view, while preserving the utility of the input data. An extensive evaluation indicates that it is possible to distribute high-quality data that respects several meaningful notions of privacy. Further, it is possible to do this efficiently for large data sets.