Fast linearization of tree kernels over large-scale data
Aliaksei Severyn, Alessandro Moschitti · 2013
Convolution tree kernels have been successfully ap-plied to many language processing tasks for achiev-ing state-of-the-art accuracy. Unfortunately, higher computational complexity of learning with kernels w.r.t. using explicit feature vectors makes them less attractive for large-scale data. In this paper, we study the latest approaches to solve such prob-lems ranging from feature hashing to reverse kernel engineering and approximate cutting plane train-ing with model compression. We derive a novel method that relies on reverse-kernel engineering together with an efficient kernel learning method. The approach gives the advantage of using tree ker-nels to automatically generate rich structured fea-ture spaces and working in the linear space where learning and testing is fast. We experimented with training sets up to 4 million examples from Seman-tic Role Labeling. The results show that (i) the choice of correct structural features is essential and (ii) we can speed-up training from weeks to less than 20 minutes. 1