Wenkai Xu’s contribution to the Discussion of ‘Safe testing’ by Grünwald, De Heide and Koolen

Wenkai Xu · Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2024

The authors propose the framework of e-value testing by investigating the grow-rate-optimal (GRO) constructions of e-variables, that incorporate a comprehensive range of testing scenarios for simple alternative, as well as their composite counterparts. The authors illustrate the versatility of e-value tests via three interpretations, drawing connections to (1) game-theoretical viewpoint, (2) conservative p-values, and (3) Bayes Factors; and devise the determination and connections for optional continuation and optional stopping with e-variables, going beyond the capacity of p-value based testing. The proposed GRO-based tests (and the corresponding variants) also embed intriguing connections with information processing, that strengthen the proposal on testing with streaming. The e-value defined by the (sub-)density in the ratio form, E=q(Y)p0(Y)⁠, naturally seeks for assessment on ratio-based measures between probabilities, i.e. ϕ-divergence Dϕ(q‖p)=∫ϕ(qp)dp, for convex function ϕ.1 The authors shed a light on the connections of e-value constructions with the Csiszár’s ϕ-divergence (f-divergence) (Csiszár, 1975) and the Reverse Information Projection (RIPr) (Csiszár, 1975; Harremoës et al., 2023). This work presents KL-based and established duality between optimal e-values and RIPr. We may further such connections through a lens of convex functions. The classical information processing inequality (Cover, 1999) reflects that the information via ϕ-divergence Dϕ(q‖p) reduces by passing through a channel (or stochastic corruption). Denote the joint distribution up to obtaining data up to n−1 by where π0 and π denotes the prior on θ in well-defined scenarios. Then, let q[n]=q(Y|θ0)q[n−1] and p0[n]=q(Y|θ0)p0[n−1]⁠. For fixed choice of ϕ, passing through the same channel q(Y|θ0) Due to the sub-density in E, E[logE]≤logE[E]≤log1⁠. Having the same channel replicates the case under the null. Dϕ(q[n−1]‖p[n−1]) in equation (1) represents the information collected to distinguish p, q with observation up to n−1⁠. The contraction of continuation under the null directly follows from data processing inequality with the same channel q(Y|θ0)⁠, which implies the ‘less’ evidence against the null and reflects the interpretation of the conservative p-value. The DPI naturally extend the analysis from KL to ϕ-divergence. The properness of log loss is essential in the analysis in the paper, while also acting as a starting motivation for the development of e-values based on general proper losses. The e-variable is designed to collect evidence for the expectation departure away from the null hypothesis under the alternative. Now let q1[n]=q(Y|θ1)q[n−1] and p0[n]=q(Y|θ0)p0[n−1]⁠, where θ1∈Θ1 and θ0∈Θ0⁠. The Jensen gap between ϕ-divergence in (1) is to be filled by Ep0[Dϕ(qθ(y)‖qθ0(y))] for the e-variable to increase, so to collect evidence. The recent development of information processing equality (Williamson & Cranko, 2024) advances the analysis of information processing, replacing the inequality by equality, where the change of measure is specified via support functions, i.e. for some corresponding choice of ϕ~⁠. The choice of ϕ and ϕ~ is then to be crafted or even optimized for interpreting evidence collection via Dϕ(qθ(y)‖qθ0(y))⁠. The Bayes factor BF:=qθ(y)p0(y) wtih the expectation form EP[BF] correspond to linear function ϕ in (1). The ϕ-divergence can be formulated in terms of Bayes risk as investigated in Reid and Williamson (2011). For some specified prior π, consider the mixture The corresponding choice of ϕ reads where L_ represents the Bayes risk with w.r.t. some choice of loss function ℓ(η,η^)2. The evidence collection procedure from E, turns the Bayes factor t=BF=QθP0 into a classification problem from mixture M in equation (3). The easier P0 and Qθ is to be distinguished, the lower the Bayes risk to achieve, the larger information Dϕ represents and the closer BF is to 1. The efficiency of the collection is then through loss function ℓ(η,η^) with two arguments corresponding to two distributions, while the GRO in-the-worst-case is based on taking inf on both hypothesis distribution classes W1∈W1(Θ1) and W0∈W0(Θ0) for KL-based measures. Through the GRO-based e-value construction, the (relative) information between data distributions and the null hypothesis can be represented in a more general fashion, embracing the Neyman, Fisher, and Jeffreys’ currency for testing, as well as generating additional insights with information processing interpretations.

Read the paper · More papers on PaperTik