The complexity of learning from a mixture of labeled and unlabeled examples

Joel Ratsaby · Scholarly Commons (University of Pennsylvania) · 1994

The learning of a pattern classification rule rests on acquiring information to constitute a decision rule that is close to the optimal Bayes rule. Among the various ways of conveying information, showing the learner examples from the different classes is an obvious approach and ubiquitous in the pattern recognition field. Basically there are two types of examples: labeled in which the learner is provided with the correct classification of the example and unlabeled in which this classification is missing. Driven by the reality that often unlabeled examples are plentiful whereas labeled examples are difficult or expensive to acquire we explore the tradeoff between labeled and unlabeled sample complexities (the number of examples required to learn to within a specified error). This problem was posed in this form by T. M. Cover and may be succinctly, if inexactly, stated as follows: How many unlabeled examples is one labeled example worth? The direction taken here focuses on the archetypal problem of learning a classification problem with two pattern classes that are typified by feature vectors, i.e., examples drawn from class conditional Gaussian distributions and where the learning approaches are parametric (Maximum Likelihood Estimation) and nonparametric (Kernel Density Estimation). Denoting the dimensionality of the example-space as N, and the number of labeled and unlabeled examples as m and n respectively, then for specific algorithms, it is shown that under a nonparametric scenario the classification error probability decreases roughly as $O((c\\sb0n\\sp{-2/N})\\sp{\\rm{lg}N})+O(e\\sp{-c\\sb1m}),$ and in the parametric scenario the error decreases roughly as $O(N\\sp{3/5}n\\sp{-1/5})+O(e\\sp{-c\\sb1m}),$ where $c\\sb0,\\ c\\sb1 >0$ are constants with respect to N, m, and n. This shows that in both scenarios exponentially more unlabeled examples are needed than labeled examples for the same reduction in error. When considering the effect of the dimensionality N, a labeled example is worth exponentially more in the nonparametric than in the parametric scenario. Extensions to the k-means procedure and Neural-Networks are also investigated.

Read the paper · More papers on PaperTik