Exploring the utility of synthetic data to extract more value from sensitive health data assets: A focused example in perinatal epidemiology
Amy Braddon, Suzanne M. Robinson, Rosa Alati, Kim Steven Betts · Paediatric and Perinatal Epidemiology · 2022
BACKGROUND: Privacy, access and security concerns can hinder the availability of health data for research. The use of synthesised data in place of de-identified electronic health records (EHRs) presents an opportunity to conduct research while minimising privacy concerns. OBJECTIVES: To examine whether synthesised data can replicate two prenatal epidemiological associations: between prenatal smoking and lower birthweight, and between prenatal mood disorders and lower birthweight, using data synthesised from de-identified health administrative data collections. METHODS: We generated two synthetic datasets, using parametric and non-parametric data generating methods, and examined the synthetic data for evidence of privacy concerns. Next, univariable and multivariable logistic regression was utilised to estimate the associations in both synthetic datasets, with results then compared to the real data. RESULTS: Both synthesised datasets performed well in identifying the reduction in birthweight associated with prenatal smoking, while the non-parametric data underestimated the reduction in birthweight associated with prenatal mood disorders. Improbable relationships between some variables were identified in the parametric synthesised data, however, these can be addressed with simple rules during data synthesis. No duplicate rows (i.e., exact copies of de-identified data) were found in the parametric data, while only 0.6% of the rows in the non-parametric data were duplicated. CONCLUSIONS: Both synthesised datasets performed well in replicating the statistical properties of the original data while addressing privacy issues. Data synthesis methods provide an opportunity for researchers to utilise health data while managing privacy and security concerns.