Model misspecification in missing data

Charles J. Geyer, Yun Ju Sung · 2003

When a statistical model is incorrect, the maximum likelihood estimate (MLE) is inconsistent, converging to the minimizer t* of Kullback-Leibler information (KLI), instead of the true parameter value. Any difference between the density ft* and the true density g is error due to model misspecification. We propose a Monte Carlo method to find t* when there are missing data and the observed data likelihood doesn't have closed form. The missing data case includes latent variables, random effects, and empirical Bayes. Our method involves generating two independent samples, the first for observed data from the true density and the second for missing data from an importance sampling density. These samples are used to approximate the KLI using importance sampling, and the minimizer of this approximation qdm,n is the Monte Carlo approximation to t* (where m and n are the missing and observed data sample sizes). We prove (Monte Carlo) consistency and asymptotic normality of qdm,n . Consistency says qdm,n can be made arbitrarily close to t* by increasing the Monte Carlo sample sizes m and n. Asymptotic normality says we can compute the Monte Carlo standard error for given m and n. This means we can estimate t* and know what accuracy we have. Also the asymptotic variance of qdm,n guides the choice of the importance sampling density. If nature, instead of a computer, generates the first sample, then our estimate is a Monte Carlo approximation to the MLE. Now its asymptotic variance reflects sampling variability of the first sample and Monte Carlo variability of the second sample. Thus our method is also a new Monte Carlo method for likelihood inference in general missing data models. Mutations, important in evolutionary biology, are studied in mutation accumulation experiments. Statistical inference for these experiments is very difficult because mutations aren't directly observable. The R package mutate has been developed to apply our method to these experiments. We discuss design and implementation of the package and its use in parallel computing.

Read the paper · More papers on PaperTik