Empirical evaluation of Gaussian Process approximation algorithms
Krzysztof Chalupka · 2011
Many large datasets are becoming available as technology advances. The fast development of the Internet makes creating huge databases of interesting data more feasible than ever; similarly, scientific simulations on modern computers can produce large amounts of information that needs to be analyzed. Machine Learning uses methods developed in Computer Science and Mathematics to deal with challenges posed in the context of such large scale data analysis. As the Bayesian framework became more popular, many flexible and theoretically elegant methods have been developed in the field. One such Bayesian framework uses Gaussian Processes to perform the two basic Machine Learning tasks, regression and classification. As it turns out, regression with Gaussian Processes is particularly elegant and analytically tractable. However, it scales badly with the size of the dataset which makes it infeasible for use in most interesting situations. Several approximation algorithms were developed to deal with this issue. While there were attempts to analyze and compare these approximations theoretically, not much has been done to present an unbiased and useful empirical evaluation of the algorithms. In this dissertation we create a solid framework for such comparison and perform experiments that allow us to analyze the practical usefulness of Gaussian Process approximation algorithms.