Discussion of the synthetic data papers published in the previous issue1
Jörg Drechsler · Statistical Journal of the IAOS · 2016
In our data driven society in which we expect that all major decisions are backed up by empirical evidence based on high quality data, broad access to these data is a must.However, the benefits of broad data access need to be balanced against potential risks of disclosure.Most data gathered by government agencies are collected under the pledge of confidentiality and the agencies have a legal and moral obligation to guarantee this pledge.Furthermore, if respondents get the impression that their data are not sufficiently protected they might refuse to participate or purposely provide wrong answers jeopardizing the quality of the collected data.Statistical agencies thus have to address this trade-off and much progress has been made in the last decades increasing the amount of data available for the general public while maintaining the confidentiality of the survey respondents.Still, there are certain types of data for which addressing this trade-off is particularly difficult.Medical records containing sensitive information on health status are one example, another example are business data.These data are particularly difficult to protect since a few variables usually suffice to identify larger businesses in the data.At the same time the collected information is often sensitive since other establishments might gain an edge if they learn certain attributes 1