Towards Privacy-Preserving Query Log Publishing
Li Ping Xiong, Eugene Agichtein · 2007
It’s an open secret that search engines collect detailed query logs, and sometimes release these data to third parties. While making this wealth of information available provides enormous opportunities for information retrieval and web mining research, it also raises serious concerns about the privacy of individuals. We strongly believe that this data should be published to allow researchers to develop new information access algorithms, however, it is desirable to anonymize these logs, so that they are still usable for research but do not contain sensitive information. The most important need is to define in a principled way the notion of privacy for query logs. This paper attempts to lay out some dimensions for defining privacy guidelines for query log publishing. We focus on the central issue of how to strike a balance between protecting the sensitive information and maintaining useful data for analysis. This work is within the overall vision of developing anonymization techniques to allow construction of IR algorithms (e.g., spelling correction) that maintain state-of-the-art performance over the anonymized data. We first describe some important applications of query log analysis and discuss their requirements on the degree of granularity of query logs. We then analyze the sensitive information in query logs and classify them from the privacy perspective. We lay out two orthogonal dimensions for anonymizing query logs and present a spectrum of approaches along those dimensions. We discuss whether existing privacy guidelines such as HIPAA can apply to query logs directly, or whether these guidelines require significant adaptation. For each of the approaches, we discuss the implications on query log utility regarding the important applications as well as the privacy of the anonymized query logs. More generally, our goal is to bring up questions and suggest challenges for privacy-preserving query log publishing.