How Differential Privacy Impacts Data Elicitation
Juba Ziani · ACM SIGecom Exchanges · 2022
Many studies and online platforms rely on data from individuals to learn properties of a population or train data-driven services and recommendation systems. However, this data is often sensitive or private, and individuals may be unwilling to share personal information. At the same time, the introduction of Differential Privacy has given us a formal and principled way to protect the privacy of individuals. Differential Privacy adds noise to the data or the learner's computation to "hide" the data of any specific individual, while still permitting to learn properties at the level of a population. In this letter, we discuss our recent work on how one may use differential privacy to increase individuals' willingness to share their personal information. We are particularly interested in understanding the following trade-off: on the one hand, providing more privacy requires adding more noise, which leads to less accurate models; on the other hand, providing more privacy allows more individuals to share their personal data, leading to better models. The main challenges are two-fold: i) we may need to provide different levels of privacy protections to individuals with differing privacy preferences and ii) we aim to incentivize individuals to share their data without payments, but instead through the utility they obtain from the data-driven service offered by the learner.