Seconder of the vote of thanks to Fong, Holmes and Walker and contribution to the Discussion of ‘Martingale Posterior Distributions’
Steffen Lilholt Lauritzen · Journal of the Royal Statistical Society Series B (Statistical Methodology) · 2023
Let me first congratulate the authors for an interesting, thought provoking and inspiring article, with many fine examples. However, I wonder whether the specific setup has sufficient generality to cover interesting cases, and I should have liked to see the ideas in this paper confronted with some different situations. Firstly, let me remind everyone of the standard Bayesian setup, where we have observablesX,Y, a parameter Θ of interest that in principle should be added to or be a function of the observables, and a Fisherian model, specifying the conditional distribution P{(X,Y)∈A×B|Θ=θ}. A Bayesian model will also specify a prior distribution π of Θ, hence also a joint distribution P{(X,Y,Θ)∈A×B×C}. Bayesian inference after observation of X=x will now calculate the posterior distribution P{Θ∈C|X=x} and/or the predictive distribution P{Y∈B|X=x}. Note in particular that this setup is universal and applies to almost any thinkable statistical problem, whereas the article specialises to a setting with X=(X1,…,Xn) and Y=Xn+1,Xn+2… (asymptotically) exchangeable, so or some variant thereof. The article then circumvents specifying prior and posterior and goes directly to the predictive. To highlight some of the issues I am thinking of, let us consider the pure birth process (Xt,t>0) specified by letting X0=1 and for t>0 To make a full Bayesian specification, we add a prior exponential distribution for the unknown parameter Λ∼exp(1). We now have the following facts, see for example Keiding (1974): Observe X on interval [0,t] and let St=∫0tXudu. Then, almost surely: Conditional on W=w, the process Xt,t>0 behaves like an inhomogeneous Poisson process (Kendall, 1949, 1966) with intensity μ(t)=wλeλt. Hence, logXt grows linearly as so W determines the intercept at 0. We now have a choice and could either think of Λ or the pair (W,Λ) as the parameter of interest. In both cases, the parameter would be a function of the data for an infinite sample size; the last parameter would give a more detailed description of the behaviour as it will not just give the slope but also the approximate intercept of (logXt,t→∞). In the first case, the log-likelihood function becomes and the MLE is λ^=(Xt−1)/St. In the second case, the log-likelihood function becomes and the MLE is more complicated and may not exist, for example if the observed growth curve is concave. In both cases, (Xt,St) is minimal sufficient. The predictive distribution for (Xt+h,St+h) given Λ=λ and Xu,u∈[0,t] has density with respect to product of counting measure and Lebesgue measure (Keiding, 1974): where gx,xt(s−st) is explicit and does not contain unknown parameters. This yields the Bayesian predictive distribution when Λ∼exp(1) as This last predictive distribution defines a ‘martingale posterior’ using that the process of sufficient statistics (Xu,Su),u>t is a Markov process. Indeed, as exploited by Doob (1949), the sequence of posterior distributions is always a martingale. Using the idea of today’s article, one could simulate from the predictive distribution and define the estimates via the simulated sample (Xu,t0; Then, the inhomogeneous Poisson model is an extreme point model (Lauritzen, 1988), but not otherwise. In any case, it seems hard to invent the predictive distribution above without going through the standard Bayesian approach so maybe the predictive approach is not so helpful after all? Is there a potential advantage in using the martingale posterior framework for describing the uncertainty of W by simulation from the predictive rather than the posterior distribution? In any case, it is my absolute pleasure to second the vote of thanks for this interesting article. The vote of thanks was passed by acclamation.