Yanbo Tang's Contribution to the Discussion of ‘Assumption-Lean Inference for Generalised Linear Model Parameters’ by Vansteelandt and Dukes
Yanbo Tang · Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2022
The method proposed by Vansteelandt and Dukes is an interesting combination of non-parametric and machine learning techniques with traditional ideas in statistics. We note in this discussion that the authors' argument is easily extended to obtain a quantitative rate in terms of the quality of the nonparametric conditional mean estimate, and comment on how this may impact a practitioner's choice of the machine learning model used to estimate the conditional mean. However the exact rate of consistency of the chosen estimator for E(Y | A, L) and the other conditional means matters if one wishes to quantify the speed at which the error term decays. For example, consider if all of the conditional means in Theorem 2 have a rate of consistency of Op {n − 1/2−ϵ}, for some ϵ > 0, then following the thread of calculation available in the appendix, we arrive at n1/2(β-β̂) = Zn + Op(n-ϵ). While the error in the approximation tends to 0, it may be doing so at a very slow rate which implies a large amount of data is needed to produce a confidence interval with the correct coverage. While if the model were a correctly specified GLM, using the traditional Wald statistic on the regression parameter β would have resulted in the classical error rate of Op (n − 1/2). This potential loss in the accuracy of the distributional approximation is not surprising as it reflects the difficulty of non-parametric estimation problems in general. This suggests that the method chosen to estimate the conditional mean needs to be flexible enough to produce a good estimator, while not so complex that the rate of consistency is too slow to be of practical use for a given dataset of size n. Thus some additional consideration is needed on the practitioner's part when selecting the machine learning or non-parametric method used to estimate the conditional mean, and in particular some knowledge of the true conditional mean function may be required, for example its smoothness in terms of higher-order derivatives.