Active Learning and Basis Selection for Kernel-Based Linear Models: A Bayesian Perspective
John William Paisley, Xuejun Liao, Lawrence Carin · IEEE Transactions on Signal Processing · 2010
We develop an active learning algorithm for kernel-based linear regression and classification. The proposed greedy algorithm employs a minimum-entropy criterion derived using a Bayesian interpretation of ridge regression. We assume access to a matrix,? ? \BBRN?N, for which the(i,j)th element is defined by the kernel functionK(?i,?j) ? \BBR, with the observed data?i? \BBRd. We seek a model,M:?i?yi, whereyiis a real-valued response or integer-valued label, which we do not have access toa priori. To achieve this goal, a submatrix,?Il,Ib? \BBRn?m, is sought that corresponds to the intersection ofnrows andmcolumns of?, indexed by the setsIlandIb, respectively. Typicallym?Nandn?N. We have two objectives:(i) Determine themcolumns of?, indexed by the setIb, that are the most informative for building a linear model,M: [1 ?i,Ib]T?yi, without any knowledge of{yi}i=1Nand(ii) using active learning, sequentially determine which subset ofnelements of{yi}i=1Nshould be acquired; both stopping values,|Ib| =mand|Il| =n, are also to be inferred from the data. These steps are taken with the goal of minimizing the uncertainty of the model parameters,x, as measured by the differential entropy of its posterior distribution. The parameter vectorx? \BBRm, as well as the model bias? ? \BBR, is then learned from the resulting problem,yIl= ?Il,Ibx+ ?1+?. The remainingN-nresponses/labels not included inyIlcan be inferred by applyingxto the remainingN-nrows of?:,Ib. We show experimental results for several regression and classification problems, and compare to other active learning methods.