Non-Mercer Large Scale Multiclass Least Squares Minimal Complexity Machines

Mayank Sharma, Sumit Soman, Jayadeva, Himanshu Pant · 2018

This paper extends the idea of Least Squares Minimal Complexity Machines [1] (LS-MCMs) to non-Mercer kernels. There are no efficient solvers for LS-SVMs with non-Mercer kernels for large scale datasets. Here, we propose two variants of our novel multiclass loss LS-MCMs. Firstly, with an L1regularizer and secondly, with an explicit margin regularizer along with L1norm in the Empirical Feature Space (EFS). Both these methods can be scaled to large datasets with the use of “prototype vectors” selected from the dataset. The first optimization algorithm can be solved efficiently using Stochastic Gradient Descent (SGD) directly as the problem remains convex in the parameter space. The second problem however, is solved using difference of convex functions (DC) programming with SGD due to the non-convex nature of margin regularizer. Our method also obtains a sparse solution as opposed to the one obtained using LS-SVMs which tend to be non-sparse.

Read the paper · More papers on PaperTik