Hierarchical Gaussian Processes for Large Scale Bayesian Regression
Marc Peter Deisenroth, A. Aldo Faisal · 2014
We present a deep hierarchical mixture-of-experts model for scalable Gaussian process (GP) regression, which allows for highly parallel and distributed computation on a large number of computational units, enabling the application of GP modelling to large data sets with tens of millions of data points without an explicit sparse representation. The key to the model is the subdivision of the data at each node into multiple child GPs performing independent computations, each with a subset of the training data. All local models share the same set of hyperparameters are trained jointly. We assume independence amongst child GPs, but this approximation is mitigated by having child GPs with overlapping subsets to account for the covariance of data in different subsets. We show that our method can match the performance of a full GP (in both prediction and hyperparameter selection by marginal likelihood maximisation) with a significantly lower computational and memory cost. We outline a number of methods for the subdivision of GPs at each hierarchy level and evaluate the robustness of each of these methods on different data sets. Finally we provide a framework for selecting the architecture the hierarchical GP for given hardware and discuss the implications of implementing our model in practice. Our practical framework allows the application of GPs to data sets with tens of millions of data points without and explicit sparse representation. Hence, our model substantially advances the current state of the art in large-scale GP modelling and finally makes it possible to apply non-parametric GPs to Big Data.