Statistical Challenges in Combining Information from Big and Small Data Sources

Trivellore Raghunathan · Deep Blue (University of Michigan) · 2015

Social Media, electronic health records, credit card transactional and administrative data, web scraping, and numerous other ways of collecting information have changed the landscape for those interested in addressing policy-relevant research questions. During the same time, the traditional sources of data, such as large-scale surveys, that have been a stable source for policy-relevant research have su ered set- backs due to large nonresponse and increasing data collection costs. The non-survey data usually contain detailed information on certain behaviors on a large number of individuals (such as all credit card transactions) but very little background information on them (such as important covariates to address the policy-relevant question). On the other hand, the survey data contains detailed information on co- variates but not so detailed information on the behaviors. Both data sources may not be perfect for the target population of interest. This paper develops and evaluates a framework for linking information from multiple imperfect data sources along with the Census data to draw statistical inference. An explicit modeling framework involving se- lection into the big data, sampling and nonresponse mechanism in the survey data, distribution of the key variables of interest and cer- tain marginal distributions from the Census Data are used as building blocks to draw inference about the population quantity of interest.

Read the paper · More papers on PaperTik