Bayesian Nonparametric Predictive Distributions (Density, Estimation, Simplex).
Peter Lenk · Deep Blue (University of Michigan) · 1984
Alternatives to the Dirichlet prior for multinomial probabilities are explored. The Dirichlet prior has the feature that the cell probabilities are only weakly correlated. A generalization of the logistic normal family on the simplex introduces a strong correlation structure between cell probabilities through the covariance of the underlying multivariate normal distribution. The posterior distribution is also a generalized logistic normal distribution where the data enter through the covariance structure. Inference about absolutely continuous distributions is analogous to the discrete case. The density is modeled by a second order, continuous in the mean, stochastic process such that neighboring ordinates are highly correlated. The infinite dimensional extension of the logistic normal distribution on the simplex is considered as a prior for the density. The logistic normal process affords three features that makes it flexible: (i) the specification of an initial density by the prior mean, (ii) the prior variance controls the influence of the data on the posterior distribution, and (iii) the prior covariance controls the amount of smoothing or pooling of the sample information. As in the discrete case, the posterior process is also a logistic normal process with the data entering through the covariance structure. The predictive density is a nonlinear function of a kernel estimator where the covariance function plays the role of the smoothing kernel. As two parameters in the covariance varies from zero to infinity, the predictive distribution systematically varies from the prior expectation to the empirical distribution, and the posterior process tends to a diffuse Dirichlet process; in which case, the predictive probability of an interval only depends on the number of observations in the interval and the sample size. A likelihood based procedure for selecting the smoothing parameter in the covariance function is proposed. The methodology is exemplified by a numerical study. In regions of low sample information the predictive density resembles the prior expectation. In regions of high sample information it combines the prior and the sample information to obtain a reasonable estimate. The prior variance determines the weight of the data, and the prior covariance controls the smoothness of the predictive distribution.