An Augmented PAC Model for Semi-Supervised Learning

Maria-Florina Balcan, Avrim L. Blum · The MIT Press eBooks · 2006

The standard PAC-learning model has proven to be a useful theoretical framework for thinking about the problem of supervised learning. However, it does not tend to capture the assumptions underlying many semi-supervised learning methods. In this chapter we describe an augmented version of the PAC model designed with semisupervised learning in mind, that can be used to help think about the problem of learning from labeled and unlabeled data and many of the different approaches taken. The model provides a unified framework for analyzing when and why unlabeled data can help, in which one can discuss both sample-complexity and algorithmic issues. Our model can be viewed as an extension of the standard PAC model, where in addition to a concept class C, one also proposes a compatibility function: a type of compatibility that one believes the target concept should have with the underlying distribution of data. For example, it could be that one believes the target should cut through a low-density region of space, or that it should be self-consistent in some way as in co-training. This belief is then explicitly represented in the model. Unlabeled data is then potentially helpful in this setting because it allows one to

Read the paper · More papers on PaperTik