Data-Driven Kernels via Semi-supervised Clustering on the Manifold
Jared Lundell, Charles DuHadway, Dan A. Ventura · 2015
We present an approach to transductive learning that employs semi-supervised clustering of all available data (both labeled and unlabeled) to produce a data-dependent SVM kernel. In the general case where the domain includes irrelevant or redundant attributes, we constrain the clustering to occur on the manifold prescribed by the data (both labeled and unlabeled). Empirical results show that the approach performs comparably to more traditional kernels while providing significant reduction in the number of support vectors used. Further, the kernel construction technique provides some of the benefits that would normally be provided by dimensionality reduction preprocessing step.