Query-based PUFs for Disclosure-Safe Remote Analysis from Medicare Claims Micro Data
A. C. Singh, Joshua Borton, Michael E. Davern, Yan‐Xia Lin · 2013
Remote analysis servers with disclosure-treatment of analysis output from user queries, as an alternative to traditional input disclosure-treated PUFs, form an active research area and are likely to be the future mode of output dissemination from data analysis. The main reason for this is that the preponderance of publically available datasets containing indirect identifiers at present or in future raises concerns about disclosure-safety of traditional PUFs based on static assumptions of intruder knowledge. However, success of remote analysis servers depends on protection from the challenging problem of possible differencing attacks by repeated queries. We describe a new application of a recently proposed method of query-based public use file (or Q-PUF) to Medicare claims data that is not vulnerable to differencing attacks. The reason for this is that in Q-PUF, the user is not allowed to arbitrarily define analysis domains but is required to choose from a checklist of pre-screened variables such that the contributor-count; i.e., the number of observations making contributions to the analysis domains defined by these variables satisfies a minimum threshold. The method is termed PUF, despite being an output treatment, because analogous to the traditional input-treated PUFs, the data producer controls the type and scope of allowed analytic variables. There are four main components of Q-PUF: first, construct a checklist of variables defining analysis domains for which the number of contributors from the data is deemed adequately large to provide reliable estimating functions of corresponding parameters, and hence rendering them automatically disclosure-safe for analysis; second, perform a disclosure audit to choose data-specific adequate confidentiality threshold; third, to provide an interface for users to communicate with the microdata via queries; and fourth, impose additional restrictions specific to the analysis output if needed. The checklist of allowed variables can be updated over time to accommodate new queries as long as disclosure-safety of existing domains is not jeopardized. We provide empirical examples as an illustration of descriptive inference using a working synthetic PUF created from Medicare claims data.