A Clustered Gaussian Process Model for Computer Experiments

Chih‐Li Sung, Benjamin Adam Haaland, Youngdeok Hwang, Siyuan Lu · Statistica Sinica · 2021

The Gaussian process has been one of the most important approaches for emulating computer simulations.However, the stationarity assumption that is common to Gaussian process emulation and computational intractability for large-scale datasets limit accuracy and feasibility in practice.In this article, we propose a clustered Gaussian process model which simultaneously segments the input data into multiple clusters and fits a Gaussian process model in each.The model parameters and the clusters are learned through the efficient stochastic expectation-maximization, which allows for emulation for large-scale computer simulations.Importantly, the proposed method provides valuable model interpretability by identifying clusters, which reveal hidden patterns in the input-output relationship.The number of clusters, which controls the bias-variance tradeoff, is efficiently selected via cross-validation to ensure accurate predictions.In our simulations as well as a real application to solar irradiance emulation, our proposed method has smaller mean squared errors than its main competitors, with competitive computation time, and provides valuable insights from data by discovering the clusters.An R package for the proposed methodology is provided in an open repository.

Read the paper · More papers on PaperTik