Decontamination of Mutually Contaminated Models
Gilles Blanchard, Clayton D. Scott · 2014
A variety of machine learning problems are characterized by data sets that are drawn from multiple different convex combinations of a fixed set of base distributions. We call this a mutual contamination model. In such problems, it is often of interest to recover these base distributions, or otherwise dis-cern their properties. This work focuses on the problem of classification with multiclass label noise, in a general setting where the noise proportions are unknown and the true class distributions are nonseparable and po-tentially quite complex. We develop a pro-cedure for decontamination of the contami-nated models from data, which then facili-tates the design of a consistent discrimina-tion rule. Our approach relies on a novel method for estimating the error when pro-jecting one distribution onto a convex combi-nation of others, where the projection is with respect to a statistical distance known as the separation distance. Under sufficient condi-tions on the amount of noise and purity of the base distributions, this projection procedure successfully recovers the underlying class dis-tributions. Connections to novelty detection, topic modeling, and other learning problems are also discussed. 1