Blair Bilodeau's Contribution to the Discussion of ‘Assumption-Lean Inference for Generalised Linear Model Parameters’ by Vansteelandt and Dukes
Blair Bilodeau · Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2022
I commend Vansteelandt and Dukes for (a) advocating that statistics move beyond well-specified assumptions and (b) providing a practical estimator to address this for generalised linear models (GLMs). Furthermore, their derivation of theoretical guarantees for this estimator is highly attractive. However, the theorem assumptions exclude many standard machine learning methods, which may violate the spirit of ‘assumption-lean’ inference guarantees. Theorems 2 and 4 assume that the three (non-parametric) conditional mean estimates have expected square loss converging at rate strictly faster than n − 1/2. Theorem 7 of Rakhlin et al. (2017) can be extended to show that if the empirical L2 entropy of the model class depends on the scale ε ∈ (0, 1) at rate ε − p for some p > 0, then the minimax expected square loss is lower bounded by n − 2/(p+2) even in the well-specified setting. That is, in order for the theory of Vansteelandt and Dukes to apply without additional assumptions, the three conditional mean models must each be strictly Donsker (p < 2), and this cannot be readily side-stepped using sample splitting as Vansteelandt and Dukes do elsewhere. The strict Donsker assumption is satisfied by models with the number of ‘parameters’ (e.g. weights of a neural network) growing strictly slower than n. However, a fixed number of parameters is exactly a parametric assumption, and in practice the number of parameters in machine learning models is often larger than n (corresponding to the interpolation regime, see Belkin et al., 2019). In the interpolation regime, such models are better understood from a non-parametric perspective; unfortunately, this means that they suffer from an exponential dependence on the dimension of the inputs. For linear models, the dimension-free entropy growth rate is ε − 2 (Mendelson & Schechtman, 2004; Zhang, 2002), and consequently, the present theory does not apply. For integer α and γ ∈ (0, 1], entropy for the class of (α, γ)-Hölder smooth functions on Rd grows at rate ε − d/(α+γ) (Wainwright, 2019, Example 5.11), requiring univariate inputs to apply the present theory in the Lipschitz (α = 0, γ = 1) setting. These examples include neural networks with linear and Lipschitz activations, as well as certain variants of random forests (Mourtada et al., 2020). In the notation of Vansteelandt and Dukes, the input dimension corresponds to the dimension of A and L jointly, and in practice L can be quite high dimensional. The authors clearly acknowledge that their convergence requirements may not be satisfied by ‘very flexible machine learning methods’. However, these requirements exclude many methods of interest, including standard neural network and random forest architectures. Ultimately, it remains an interesting open question whether Vansteelandt and Dukes' estimator enjoys convergence guarantees (even with appropriately slow rates) when used with such canonical nonparametric methods.