Training data analysis for Gaussian process state space models

Patricia Ferreiro Alonso · 2017

Gaussian Process State Space Models aim at constructing models of nonlinear dynamical systems capable of quantifying the uncertainty in their predictions. By means of sampling in a noisy environment and covariance functions, Gaussian Process regression techniques aim to infer an estimate of the underlying function as well as a probabilistic confidence interval. Optimally choosing sample points is crucial for system identification and control as it conforms, together with the prior knowledge, all the information available to approach the inference problem. The error between the real system and the estimation, as well as its probabilistic confidence interval, directly depend on a measure of the true function complexity, the maximum information gain and the number and distribution of the training data for a given kernel. However, a closed-form solution for the aforementioned parameters hasn't been presented in past literature. In this work, we show proof of exact information confidence bounds for the Linear Kernel and derive a connection between its parameters and the most informative subset of sample points. We derive closed forms for the information maximization problem, thus avoiding a non-linear optimization problem and significantly reducing the computational load. We also compute the true function's norm in its associated Reproducing Kernel Hilbert Space and use it as a measure of complexity of the true function. Finally, we obtain a unique sample point distribution that ensures both minimal sample variance and maximum information gain for the Linear Kernel. Additionally, a similar intuition is developed for the Gaussian Kernel, computing the true function norm in terms of its Fourier transform and deriving a similar connection between the sample point distribution and the tightest confidence bounds.

Read the paper · More papers on PaperTik