Storing and querying data securely in untrusted environments
Sharad Mehrotra, Bijit Hore · 2007
In recent times, a lot of interest has been generated in understanding the nature of information disclosure in a variety of data-centric applications, especially the kind that may lead to privacy violation. In this thesis, we describe our work in context of following three applications where these issues assume prime importance: (i) Database-As-a-Service (DAS). The goal of such a system is to enable individuals (or organizations) to store data securely on a remote service-provider's server and allow them to query/modify it at any later time as needed. The two main objectives of DAS are ensuring security of client's data at all times, and offering a useful set of server-side query processing capabilities. In scenarios where the trust placed on the service-provider is limited and the data needs to be kept encrypted, satisfying these two objectives simultaneously becomes a challenge. We develop data-partitioning based schemes to support an important class of SQL queries in this model. Further, we formalize the notion of disclosure-risk and develop trade-off algorithms to optimally balance the two competing goals of performance and security in such a system. (ii) Pervasive Spaces. We design a system for privacy-preserving event detection in a pervasive environment. The system captures and stores individual-centric information in order to detect complex of interest. We use surveillance as a motivating application where the goal is to detect events related to access-violations and suspicious activity in a monitored pervasive space. The privacy goal is that of ensuring a specified level of at all times, for every individual associated with the space unless a rule is violated, which requires the identity to be disclosed. We formalize the notion of anonymity and precisely characterize the nature of inference channels that exist in such a system. We illustrate the difficulty of ensuring k-anonymity in the general case and propose a class of solutions for the same. (iii) Privacy-preserving data publishing. In this application we studied the problem of data publishing for mining applications, where the goal is to sanitize individual centric data sets for publishing so that no specific individual can be identified in the released data set. We devise a novel, enumeration-based hierarchical data partitioning scheme that allows one to determine suitable generalization of a multidimensional data set and carry out extensive experimentations to confirm the superiority of our algorithms.