Learning from partially labeled data

Marcin Szummer · 2002

The Problem: Learning from data with both labeled training points (x,y pairs) and unlabeled training points (x alone). For the labeled points, supervised learning techniques apply, but theycannot take advantage of the unlabeled points. On the other hand, unsupervised techniques can model the unlabeled data distribution, but do not exploit the labels. Thus, this task falls between traditional supervised and unsupervised learning. Motivation: Supervised learning performance improves with larger training data sets. Unfortunately, it is often infeasible to obtain labels for large training sets. Assigning labels can require expensive resources such as human labor or laboratory tests. In some cases ground truth labels are impossible to obtain, e.g. if the necessary measurements can no longer be made, or if the labels will be given only in the future. In contrast, unlabeled training data is frequently easy to obtain in large quantities, and can outnumber the amount of labeled data by a large factor. For example, it is expensive to collect image databases of only faces, but it is cheap to collect arbitrary imagery with occasional faces, e.g. by crawling the world wide web, or by pointing a video camera out the window. There are also developmental motivations for studying the process of learning from partially data. Children acquire language mainly by listening and imitating, with verylimited feedback from adults. Human beings also excel at other partially labeled learning tasks, suchasvisualdiscrimination with hyperacuity [2]. Previous Work: Learning from partially labeled data is not well understood from the theoretical perspective. Labeled

Read the paper · More papers on PaperTik