Using the Equivalent Kernel to Understand Gaussian Process Regression

Peter Sollich, Christopher K. I. Williams · 2004

The equivalent kernel [1] is a way of understanding how Gaussian process regression works for large sample sizes based on a continuum limit. In this paper we show (1) how to approximate the equivalent kernel of the widely-used squared exponential (or Gaussian) kernel and related kernels, and (2) how analysis using the equivalent kernel helps to understand the learning curves for Gaussian processes. Consider the supervised regression problem for a dataset D with entries (xi, yi) for i = 1,..., n. Under Gaussian Process (GP) assumptions the predictive mean at a test point x∗ is given by ¯f(x∗) = k ⊤ (x∗)(K + σ 2 I) −1 y, (1) where K denotes the n × n matrix of covariances between the training points with entries k(xi, xj), k(x∗) is the vector of covariances k(xi, x∗), σ 2 is the noise variance on the observations and y is a n × 1 vector holding the training targets. See e.g. [2] for further details.

Read the paper · More papers on PaperTik