Rachael V. Phillips and Mark J. van der Laan’s Contribution to the Discussion of ‘Assumption-Lean Inference for Generalised Linear Model Parameters’ by Vansteelandt and Dukes
Rachael V. Phillips, Mark J. van der Laan · Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2022
We commend the authors on this novel non-parametric extension of main and interaction coefficients in a generalised linear model (GLM), such as E(Y|A,L)=βA+g(L), and their efficient estimation. There is a wealth of non-parametric extensions one could pursue, including weighted squared error projections of a conditional treatment effect E(Y|A=a,L)-E(Y|A=0,L) and conditional interaction effect E(Y|A=(a1+1,a2+1),L)-E(Y|A=(a1,a2+1),L)-E(Y|A=(a1+1,a2),L)+E(Y|A=(a1,a2),L) onto a simple parametric models βa and β0+β1a1+β2a2, respectively (Chambaz et al. (2012)). Such projection estimands are common in the causal inference literature. The authors propose a different estimand criterion that requires the efficient influence curve (EIC) to avoid (conditional) density estimation and inverse weighting. Inverse weighting can certainly cause instability but machine learning techniques have been well-adapted for conditional density estimation (for instance, Muñoz and van der Laan (2011), Rytgaard et al. (2021) and van der Laan (2010)). The above least squares projection for the main effect has an EIC that inverse weights by P(A=0|L), thereby achieving the first but not the second goal. The proposed main effect estimand E((A-E(A|L))(E(Y|A,L)-E(Y|L)))/EσA|L2 (5) satisfies this criteria for the EIC; however, it is harder to interpret outside the GLM. For example, if (E(Y|A,L)-E(Y|L)) is not linear in (A-E(A|L)), then the numerator will generally average both negative and positive contributions E(Y|A,L)-E(Y|L), even for problems in which E(Y|A,L)-E(Y|A=0,L)≥0 everywhere. Consider A∼U(0,1) independent of L and E(Y|A,L)=ϵ-1AI(A≤ϵ)+I(A>ϵ)+g(L). Here, the numerator becomes E(A-1/2)(ϵ-1AI(A≤ϵ)+I(A>ϵ)) and approximates E(A-1/2)=0 as ϵ→0, whereas the unweighted projection of E(Y|A,L)-E(Y|A=0,L) on βa, would result in β=3/2-ϵ2/2, correctly demonstrating a strong treatment effect. In addition to impacting the interpretation, these cancellations can hurt the power for testing a null hypothesis relative to testing with the parameter in Chambaz et al. (2012), even though the latter’s EIC might have larger variance. The interpretation of the interaction estimand (e.g., 10) for continuous A1, A2 has an additional complication in the sense that it does not involve an average of L-specific interactions. Both of the proposed estimands and their EICs depend on the conditional distribution of A, given L, making the interpretations non-robust to irrelevant deviations in the study. Users should carefully consider knowledge regarding the model for P(A|L) and compute the EIC accordingly, potentially achieving efficiency gains relative to the non-parametric model considered. Finally, the authors proposed a one-step estimator of their estimands based on the EIC. We believe that practitioners that are used to maximum likelihood estimation (MLE) would find it more natural to use a plug-in estimator, and thereby targeted MLE (TMLE). It would be straightforward to develop a TMLE, as the EICs imply how to target initial estimators of E(A|L) and E(Y|A,L) for the main terms estimand, and their analogs for the interaction estimand.