On privacy-preserving data publishing and analysis
Jeffrey F. Naughton, Yeye He · 2012
Data have exploded in recent years. The availability of data at ever finer granularity presents immense potential for uncovering knowledge and producing economic value in a variety of domains. However, much of these interesting data are private. Owners of such data have legal and ethical responsibilities to protect the privacy of individuals represented by the data. Perhaps for the fear of inadvertently breaching privacy, most private data are locked-up by data owners, hampering the data from being fully utilized. Developing privacy-preserving techniques to analyze data without compromising privacy is thus an important direction of research in the field of data management. This dissertation explores privacy-preserving data anonymization techniques in three different data models. It first studies the problem of anonymizing relational data. We observe that current state-of-the-art approach still leaks sensitive information when relational data are updated. We propose techniques to address this without significantly sacrificing utility. Set-valued data is another important data model, examples of which include market basket data and query logs. In the second part of this dissertation we propose a new top-down algorithm to anonymize set-valued data that achieves better utility than previous algorithms. Lastly, we focus on the popular streaming data model, which has attracted little attention for privacy research so far. We identify potential privacy concerns of a special form of stream processing called Complex Event Processing, and study the corresponding optimization problem of maximizing utility under privacy constraints in the last part of this dissertation.