Role of Majorization in Learning the Kernel within a Gaussian Process Regression Framework
Prasenjit Kapat · OhioLink ETD Center (Ohio Library and Information Network) · 2011
Over the recent years, machine learning techniques have breathed a new life in to the classical regression framework.The primary focus in these techniques has often been the predictive performance of the estimated models and the models themselves have developed in to sophisticated non-linear predictive machines.In this development, the ubiquitous "kernel-trick" has played a very important role by providing a means to compute the inner products in the unwieldy high-dimensional spaces via simple and easily computable functions on the low-dimensional covariate domains, called as kernels.The domain knowledge of data dictates the collection of kernels 1 suitable for the specific application.In "learning the kernel" paradigm, current state of the art is to use some optimization method to select the best kernel for the data at hand from this collection.The work in this dissertation assumes the existence of a "true" underlying process, a Gaussian Process, (defined by a fully specified covariance kernel) for the given data.The Gaussian Process itself is considered as a prior on the reproducing kernel Hilbert space of functions characterized by the associated kernel.The goal is to make suggestions towards developing some diagnostic tools which can be used to hasten the kernel learning process.In particular, the setup for computational experimentation xviii A.38 Functional norm (2.2.2) at λ = λ 0 : log ν (h, λ 0 ) as a function of log(h) using Y tr 9 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2h 0 .Dotted line is at log(h 0 ) (vertical) and log(n) ≈ 5.3 (horizontal). . . .120 A.39 Functional norm (2.2.2) at λ = λ 0 : log ν (h, λ 0 ) as a function of log(h) using Y tr 10 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2h 0 .Dotted line is at log(h 0 ) (vertical) and log(n) ≈ 5.3 (horizontal). . . .121 B.1 Comparing GCV(h, λ 0 ) (solid blue curve) and Eblue curve) and E h,λ 0 GCV(h, λ 0 ) (dashed red curve) as a function of log(h) using Y tr 5 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2h 0 .Dotted line is at log(h 0 ). . . . . . . . .124 B.5 Comparing GCV(h, λ 0 ) (solid blue curve) and E h,λ 0 GCV(h, λ 0 ) (dashed red curve) as a function of log(h) using Y tr 6 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2h 0 .Dotted line is at log(h 0 ). . . . . . . . .124 B.6 Comparing GCV(h, λ 0 ) (solid blue curve) and E h,λ 0 GCV(h, λ 0 ) (dashed red curve) as a function of log(h) using Y tr 7 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2h 0 .Dotted line is at log(h 0 ). . . . . . . . .125 B.7 Comparing GCV(h, λ 0 ) (solid blue curve) and E h,λ 0 GCV(h, λ 0 ) (dashed red curve) as a function of log(h) using Y tr 8 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2h 0 .Dotted line is at log(h 0 ). . . . . . . . .125 B.8 Comparing GCV(h, λ 0 ) (solid blue curve) and E h,λ 0 GCV(h, λ 0 ) (dashed red curve) as a function of log(h) using Y tr 9 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2h 0 .Dotted line is at log(h 0 ). . . . . . . . .126 xix C.27 Diagnostic tool based on GCV (2.4.1):Γ (h, λ 0 ) as a function of log(h) using Y tr 10 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2 h 0 .Dotted lines at log(h 0 ) (vertical) and 0 (horizontal). . . . . . . . . . .146 C.28 Diagnostic tool based on functional norm (2.4.2):Υ (h, λ 0 ) as a function of log(h) using Y tr 2 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2 h 0 .Dotted lines at log(h 0 ) (vertical) and 0 (horizontal). . . . . . . .147 C.29 Diagnostic tool based on functional norm (2.4.2):Υ (h, λ 0 ) as a function of log(h) using Y tr 3 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2 h 0 .Dotted lines at log(h 0 ) (vertical) and 0 (horizontal). . . . . . . .148 C.30 Diagnostic tool based on functional norm (2.4.2):Υ (h, λ 0 ) as a function of log(h) using Y tr 4 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2 h 0 .Dotted lines at log(h 0 ) (vertical) and 0 (horizontal). . . . . . . .148 C.31 Diagnostic tool based on functional norm (2.4.2):Υ (h, λ 0 ) as a function of log(h) using Y tr 5 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2 h 0 .Dotted lines at log(h 0 ) (vertical) and 0 (horizontal). . . . . . . .149 C.32 Diagnostic tool based on functional norm (2.4.2):Υ (h, λ 0 ) as a function of log(h) using Y tr 6 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2 h 0 .Dotted lines at log(h 0 ) (vertical) and 0 (horizontal). . . . . . . .149 C.33 Diagnostic tool based on functional norm (2.4.2):Υ (h, λ 0 ) as a function of log(h) using Y tr 7 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2 h 0 .Dotted lines at log(h 0 ) (vertical) and 0 (horizontal). . . . . . . .150 C.34 Diagnostic tool based on functional norm (2.4.2):Υ (h, λ 0 ) as a function of log(h) using Y tr 8 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2 h 0 .Dotted lines at log(h 0 ) (vertical) and 0 (horizontal). . . . . . . .150 C.35 Diagnostic tool based on functional norm (2.4.2):Υ (h, λ 0 ) as a function of log(h) using Y tr 9 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2 h 0 .Dotted lines at log(h 0 ) (vertical) and 0 (horizontal). . . . . . . .151 C.36 Diagnostic tool based on functional norm (2.4.2):Υ (h, λ 0 ) as a function of log(h) using Y tr 10 .Top row: 1.5 ≤ h ≤ 60; bottom row: h 0 /2 ≤ h ≤ 2 h 0 .Dotted lines at log(h 0 ) (vertical) and 0 (horizontal). . . . . . . .151 xxiv Chapter 1: RKHS Theory, examples in Statistical Learning, Gaussian Processes, Hadamard Products, and MajorizationOver the past couple of decades, predictive learning algorithms have been the focus of a lot of research both in the field of Statistics ("statistical learning") and Computer Science ("data mining").The remarkable success of these algorithms were fuelled by the growing computational capabilities and the mathematical theory of Reproducing Kernel Hilbert Spaces (RKHS) put forth by Aronszajn (1943Aronszajn ( , 1950) ) and introduced to the Statistical community by Kimeldorf and Wahba (1971).The current chapter describes the underlying setup for "learning the kernel" problem and states some of the necessary results from literature that are needed for the current work.It is structured into a few sections.introducing the RKHS theory, giving some examples from the statistical literature where this framework is used, and the various types of kernels used in practice.It also briefly describes the setup for Gaussian Processes, the framework used in this research, and its connection to reproducing kernel Hilbert spaces.Following this, a little discussion on the kernel estimation literature in general and finally some results on Hadamard product of matrices and majorization of vectors is introduced.