Gaussian Process Regression with Heteroscedastic Residuals and Fast MCMC Methods

Chun-Yi Wang · TSpace (University of Toronto) · 2014

Gaussian Process (GP) regression models typically assume that residuals are Gaussian and have the same variance for all observations. However, applications with input-dependent noise (heteroscedastic residuals) frequently arise in practice, as do applications in which the residuals do not have a Gaussian distribution. In this thesis, we propose a GP regression model with a latent variable that serves as an additional unobserved covariate for the regression. This model (which we call GPLC) allows for heteroscedasticity since it allows the function to have a changing partial derivative with respect to this unobserved covariate. With a suitable covariance function, our GPLC model can handle (a) Gaussian residuals with input-dependent variance, or (b) non-Gaussian residuals with input-dependent variance, or (c) Gaussian residuals with constant variance. We compare our model, using synthetic datasets, with a model proposed by Goldberg, Williams and Bishop (1998), which we refer to as GPLV, which only deals with case (a), as well as a standard GP model which can handle only case (c). Markov Chain Monte Carlo methods are developed for the GPLC and GPLV models. Experiments show that when the data isheteroscedastic, both GPLC and GPLV give better results (smaller mean squared error and smaller negative log-probability density) than standard GP regression. In addition, if we do not assume Gaussian residuals, our GPLC model (as in case (b) above) is still generally nearly as good as GPLV when the residuals are in fact Gaussian. When the residuals are non-Gaussian, our GPLC model is better than GPLV.Evaluating the posterior probability density function is the most costly operation when Markov Chain Monte Carlo (MCMC) is applied to many Bayesian inference problems. For GP models, computing the posterior density involves computing the covariance matrix, and then inverting the covariance matrix. The computation time for computingthe covariance matrix is proportional to pn2, and for the inversion is proportional to n3, where p is the number of covariates and n is the number of training cases. We introduce MCMC methods based on the "temporary mapping and caching" framework (Neal, 2006), using a fast approximation, *, as the distribution needed to construct thetemporary space. We propose two implementations under this scheme: "mapping to a discretizing chain", and "mapping with tempered transitions", both of which are exactly correct MCMC methods for sampling , even though their transitions are constructed using an approximation. These methods are equivalent when their tuning parameters are set at the simplest values, but differ in general. We compare how well these methods work when using several approximations, nding on synthetic datasets that a * based on the "Subset of Data" (SOD) method is almost always more efficient than standard MCMC using only On some datasets, a more sophisticated * based on the "Nystrom-Cholesky" method works better than SOD.

Read the paper · More papers on PaperTik