Controlling the Bias-Variance Tradeoff via Coherent Risk for Robust Learning with Kernels
Alec Koppel, Amrit Singh Bedi, Ketan Rajawat · 2019
In supervised learning, we learn a statistical model by minimizing a measure of fitness averaged over data. Doing so, however, ignores the variance, i.e., the gap between the optimal within a hypothesized function class and the Bayes Risk. We propose to account for both the bias and variance by modifying training to incorporate coherent risk which quantifies the uncertainty of a given decision. We develop the first online solution to this problem when estimators belong to a reproducing kernel Hilbert space (RKHS), which we call Compositional Online Learning with Kernels (COLK). COLK addresses the fact that (i) minimizing risk functions requires handling objectives which are compositions of expected value problems by generalizing the two time-scale stochastic quasi-gradient method to RKHSs; and (ii) the RKHS-induced parameterization has complexity which is proportional to the iteration index which is mitigated through greedily constructed subspace projections. We establish linear convergence in mean to a neighborhood with constant stepsizes, as well as the fact that its complexity is at-worst finite. Experiments on synthetic and benchmark data demonstrate that COLK exhibits consistent performance across training runs, estimates that are both low bias and low variance, and thus marks a step towards overcoming overfitting.