A Gaussian Latent Variable Model for Large Margin Classification of Labeled and Unlabeled Data

Do-kyum Kim, Matthew F. Der, Lawrence K. Saul · 2014

We investigate a Gaussian latent variable model for semi-supervised learning of linear large mar-gin classifiers. The model’s latent variables en-code the signed distance of examples to the sep-arating hyperplane, and we constrain these vari-ables, for both labeled and unlabeled examples, to ensure that the classes are separated by a large margin. Our approach is based on simi-lar intuitions as semi-supervised support vector machines (S3VMs), but these intuitions are for-malized in a probabilistic framework. Within this framework we are able to derive an es-pecially simple Expectation-Maximization (EM) algorithm for learning. The algorithm alternates between applying Bayes rule to “fill in ” the la-tent variables (the E-step) and performing an un-constrained least-squares regression to update the weight vector (the M-step). For the best results it is necessary to constrain the unlabeled data to have a similar ratio of positive to negative exam-ples as the labeled data. Within our model this constraint renders exact inference intractable, but we show that a Lyapunov central limit theorem (for sums of independent, but non-identical ran-dom variables) provides an excellent approxima-tion to the true posterior distribution. We perform experiments on large-scale text classification and find that our model significantly outperforms ex-isting implementations of S3VMs. 1

Read the paper · More papers on PaperTik