David Huk, Lorenzo Pacchiardi, Ritabrata Dutta and Mark Steel's contribution to the Discussion of ‘Martingale posterior distributions’ by Fong, Holmes and Walker
David Huk, Lorenzo Pacchiardi, Ritabrata Dutta, Mark F. J. Steel · Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2023
We congratulate the authors for this thought-provoking paper. The recursive definition of the predictive distributions through a bivariate copula (Section 4) depends on a hyperparameter ρ, which the authors tune by minimising the prequential log-score −∑i=1nlogpi−1(yi) (Section 4.5.2). The prequential framework nicely connects with the predictive resampling approach used later. However, other strictly proper scoring rules (Gneiting & Raftery, 2007) could be used in place of the log-score, leading to a generic prequential score ∑i=1nS(Pi−1,yi), where S(Pi−1,yi) is a scoring rule between the distribution Pi−1 and data yi. With exchangeable data, the prequential log-score is the only prequential score invariant to permutations of y1:n (as it corresponds to the log marginal −logp1:n(y1:n), Fong & Holmes, 2020) and is thus a natural choice over other scoring rules. Nevertheless, in the present set-up exchangeability is forsaken by defining the predictive distributions directly (as the authors remark in Section 3.2 and address in Section 4.5.3); indeed, computing the prequential log-score on multiple permutations of the data leads to different values, as each pi is defined iteratively from pi−1. Therefore, there seems to be no theoretical reason to prefer the log-score over other strictly proper scoring rules. We also believe the connection to cross-validation mentioned by the authors relies on exchangeability of the data through the invariance of the marginal likelihood to data permutations (Fong & Holmes, 2020). Besides, while the form of predictive distribution used by the authors provides access to the density and thus enables convenient computation of the log-score, extensions of this work could rely on predictive distributions whose density can be computed only up to a normalising constant. Furthermore, one could employ predictive distributions for which simulation is possible but density evaluation is not (relying, for instance, on generative neural networks). In both these cases, the log-score would be inaccessible, while other scoring rules [the Hyvärinen score (Hyvärinen, 2005) in the former case and the energy or kernel score (Gneiting & Raftery, 2007) in the latter] would enable hyperparameter tuning. Interestingly, the kernel score enjoys robustness to outliers in the data in different set-ups (Chérief-Abdellatif & Alquier, 2022; Pacchiardi & Dutta, 2021); although we are unsure if this property translates to the considered framework, this is worth investigating. As a first test, we tuned ρ for the univariate Gaussian mixture model in Section 5.1.1 with the prequential energy score (estimated using an importance sampling strategy) and obtained comparable values of ρ to the ones reported by the authors with the log-score.