Sequential Learning with LS-SVM for Large-Scale Data Sets
Tobias Jung, Daniel Polani · 2006
Abstract. We present a subspace-based variant of LS-SVMs (i.e. reg-ularization networks) that sequentially processes the data and is hence especially suited for online learning tasks. The algorithm works by se-lecting from the data set a small subset of basis functions that is sub-sequently used to approximate the full kernel on arbitrary points. This subset is identied online from the data stream. We improve upon ex-isting approaches (esp. the kernel recursive least squares algorithm) by proposing a new, supervised criterion for the selection of the relevant basis functions that takes into account the approximation error incurred from approximating the kernel as well as the reduction of the cost in the original learning task. We use the large-scale data set 'forest ' to compare performance and eciency of our algorithm with greedy batch selection of the basis functions via orthogonal least squares. Using the same num-ber of basis functions we achieve comparable error rates at much lower costs (CPU-time and memory wise).