The Effect of Missing Values for Covariates
Gary George Whitlock, Taane Gregory Clark · Epidemiology · 2002
To the Editor: Several investigators have recommended the change-in-estimate method of selecting covariates, 1–3 but we suggest that this method should be used with special caution if the covariate in question has a large proportion of missing values. Most analysis programs will automatically delete records that have missing values of covariates, and these deletions may in themselves cause important changes in effect estimates, whether or not the covariate is a true confounder. We illustrate this point below using data from a cohort study of 10,525 New Zealand men and women. 4 In this analysis, we investigated the potential for neighborhood income to confound the relation between handedness and risk of motor vehicle driver injury. Using the change-in-estimate approach, when neighborhood income was added to a base model that had covariates for handedness, age, sex, driving exposure, alcohol intake, and rural residence status, the hazard ratio for left-handedness increased from 1.32 (95% CI = 0.79–2.21) to 1.60 (95% CI = 0.92–2.78) (reference group: right-handedness). This 21% change in the point estimate thereby implied that neighborhood income may have been a moderately important confounder. 1 However, this conclusion was at odds with a lack of evidence that neighborhood income was independently associated with either the exposure or the outcome. The age- and sex-standardized prevalence of left-handedness across quartiles of neighborhood income ranged merely from 9.6–10.1%. Furthermore, in the quartile with the highest driver injury risk, driver injury was only 1.63 times (95% CI = 0.80–3.31) as likely as in the quartile with the lowest risk (adjusted for the same covariates as above). This large change in the effect estimate appears to have been caused chiefly by deletion of records with missing values for neighborhood income (11% of participants), rather than by the adjustment procedure itself. Using an approach described elsewhere, 5 the hazard ratio obtained from the base model after records with missing values for neighborhood income had been deleted (and without including neighborhood income as a covariate) was also 1.60 (95% CI = 0.92–2.78) –ie, the same, to two decimal places, as after adjustment for neighborhood income. It consequently seems safe to assume that neighborhood income caused little, if any, confounding, and that participants with missing values for neighborhood income were in some way unrepresentative of those analyzed in the base model. Notably, the driver injury incidence rate for participants with missing values of neighborhood income (24.9 per 10,000 person-years) was about twice as high as that for other participants (11.9 per 10,000 person-years). It has been reported previously that large changes in effect estimates upon addition of covariates to regression models do not necessarily indicate confounding by the new variables. Other possible causes include adjustment for collinear variables, intermediate variables, variables measured with error, or too many covariates for the particular fitting procedure (the so-called finite sample bias). 1,2,6 Adjustment for covariates with missing values can be added to this list. Gary Whitlock Taane Clark