Risks of Semi-Supervised Learning: How Unlabeled Data Can Degrade Performance of Generative Classifiers

Fábio Gagliardi Cozman, Ira L. Cohen · The MIT Press eBooks · 2006

Empirical and theoretical results have often testified favorably to the semisupervised learning of generative classifiers, as described in other chapters of this book.However, the literature has also brought to light a number of situations where semi-supervised learning fails to produce good generative classifiers.Here some clarification is due.We are not simply concerned with classifiers that produce high classification error -this can also happen in supervised learning.Our concern is this: it is frequently the case that we would be better off just discarding the unlabeled data and employing a supervised method, rather than taking a semi-supervised route.Thus we worry about the embarrassing situation where the addition of unlabeled data degrades the performance of a classifier.How can this be?Typically we do not expect to be better off by discarding data; how can we understand this aspect of semi-supervised learning?In this chapter we focus on the effect of modeling errors in semi-supervised learning, and show how modeling errors can lead to performance degradation.This is a portion of the eBook at doi:10.

Read the paper · More papers on PaperTik