New Approaches to Data Dissemination: A Glimpse into the Future (?)
Jerome P. Reiter · CHANCE · 2004
Many national statistical agencies and survey organizations disseminate microdata, i.e., data on individual units in public use data files. These data disseminators strive to release files that are safe from attacks by illintentioned data users seeking to learn respondents’ identities or sensitive attributes, informative for a wide range of statistical analyses, and easy for users to analyze with standard statistical methods. Meeting all three goals is a challenging task. The proliferation of readily available databases, and advances in statistical and computing technologies, provide users with more and higher quality resources for linking records in released datasets to units in other databases. As a result, the risk of unintended or illegal disclosures is high and still rising. Microdata proliferation and statistical advances also enable and fuel the ambition of researchers. To address complex statistical questions, these users demand greater access to accurate data at fine levels of detail. Data disseminators thus find themselves in a difficult position: users pressure them to provide everything about the data, but disclosure risks pressure them to limit what is released. Data disseminators that fail to prevent disclosures of individuals’ identities or sensitive attributes can face serious consequences. They may be in violation of laws and therefore subject to legal actions; they may lose the trust of the public, so that respondents are less willing to parWhat roles do remote servers and synthetic data play in the future of data dissemination?