Privacy Preserving Data Publishing for Recommender System

Xiaoqiang Chen, Vincent Huang · 2012

Driven by mutual benefits, exchange and publication of data among various parties is an inevitable trend. However, released data often contains sensitive user information thus direct publication violates individual privacy. Among many privacy models, k-anonymity framework is popular and well-studied, it protects information by constructing groups of anonymous records such that each record in the table released is covered by no fewer than k-1 other records. In this paper, we first investigate different privacy preserving technologies and then focus on achieving k-anonymity for large scale and sparse databases, especially recommender systems. We present a general process for anonymization of large scale database. A preprocessing phase strategically extracts preference matrix from original data by Singular Value Decomposition (SVD) and eliminates the high dimensionality and sparsity problem. We developed a new clustering based k-anonymity heuristic named Bisecting K-Gather (BKG) and it is proven to be efficient and accurate. To support customized user privacy assignments, we also proposed a new concept called customized k-anonymity along with a corresponding algorithm (BOKG). We use MovieLens database to assess our algorithms. The results show that we can efficiently release anonymized data without compromising the utility of data.

Read the paper · More papers on PaperTik