Causal Models
Beverly J. Levine · Epidemiology · 2009
To the Editor: George Orwell wrote that language could be used to give the “… appearance of solidity to pure wind.”1 It is disturbing that the language of “causal modeling” is being used to bestow the solidity of the complex process of causal inference upon mere statistical analysis of observational data. Hogan2 implies that because we think we are estimating causal parameters with marginal structural models, we should call these “causal models.” A cardinal rule of good science is to not accept as truth everything we think. By Hogan's implication, one who is convinced that he is estimating a causal effect in observational data with a simple regression model, or even a 2 × 2 table, should say he is using a “causal model.” Because marginal structural models can better address time-dependent confounding, they may provide less biased causal effect estimates than standard methods. However, the assumptions underlying such models are stringent, and the reality is that, as Hogan notes (p.432), “… causal models cannot be fully identified from observational data; that is, the parameters of interest cannot be estimated without making untestable assumptions.” Thus, it is not obvious that these models will improve our ability to draw accurate causal conclusions. History suggests there is no positive correlation between statistical complexity and accuracy of causal inference. One can test whether marginal structural models are better only if results from experiments are available against which to compare already-obtained results from observational data. (Finding, posthoc, that one can reproduce findings from an experiment with a marginal structural model does not necessarily speak to superiority of these models over other methods.) For many questions, then, we may never know whether marginal structural models promote accurate causal inference. The issue of terminology is especially important if, as Hogan presages, the reliance on large observational data sets will only increase. To promote language that will assuredly encourage many who are not statisticians to believe they understand a complex causal relation because they've used a complicated model on observational data—ie, a “causal model”—is a disservice that scholars should avoid. The first section of Hogan's commentary is titled “Modify the language.” Let's start with the misleading term of “causal model,” and stop conflating statistical and causal inference. The phrases “marginal structural model,” “inverse probability weighted model,” and “model accounting for time-dependent confounding” are accurately descriptive, and do not over-promise. Orwell would have urged us to refrain from letting our thought corrupt our language, and from letting our language corrupt others’ thought. Beverly Levine Department of Public Health Education, University of North Carolina, Greensboro, NC [email protected]