Kernel Methods and Support Vector Machines
Bernhard Schölkopf, Alex J. Smola · 2003
Introduction Over the past ten years kernel methods such as Support Vector Machines and Gaussian Processes have become a staple for modern statistical estimation and machine learning. The groundwork for this field was laid in the second half of the 20th century by Vapnik and Chervonenkis (geometrical formulation of an optimal separating hyperplane, capacity measures for margin classifiers), Mangasarian (linear separation by a convex function class), Aronszajn (Reproducing Kernel Hilbert Spaces), Aizerman, Braverman, and Rozonoer (nonlinearity via kernel feature spaces), Arsenin and Tikhonov (regularization and ill-posed problems), and Wahba (regularization in Reproducing Kernel Hilbert Spaces). However, it took until the early 90s until positive definite kernels became a popular and viable means of estimation. Firstly this was due to the lack of su#ciently powerful hardware, since kernel methods require the computation of the socalled kernel matrix, which requires quadratic storage i