An information-theoretic framework for learning models of instance-independent label noise
Xia Huang, Kai Fong Ernest Chong · 2021
Given a dataset D with label noise, how do we learn its underlying noise model? If we assume that the label noise is instance-independent, then the noise model can be represented by a noise transition matrix QD. Recent work has shown that even without further information about any instances with correct labels, or further assumptions on the distribution of the label noise, it is still possible to estimate QD while simultaneously learning a classifier from D. However, this presupposes that a good estimate of QD requires an accurate classifier. In this paper, we show that high classification accuracy is actually not required for estimating QD well. We shall introduce an information-theoretic-based framework for estimating QD solely from D (without additional information or assumptions). At the heart of our framework is a discriminator that predicts whether an input dataset has maximum Shannon entropy, which shall be used on multiple new datasets D^ synthesized from D via the insertion of additional label noise. We prove that our estimator for QD is statistically consistent, in terms of dataset size, and the number of intermediate datasets D^ synthesized from D. As a concrete realization of our framework, we shall incorporate local intrinsic dimensionality (LID) into the discriminator, and we show experimentally that with our LID-based discriminator, the estimation error for QD can be significantly reduced. We achieved average Kullback--Leibler loss reduction from 0.27 to 0.17 for 40% anchor-like samples removal when evaluated on the CIFAR10 with symmetric noise. Although no clean subset of D is required for our framework to work, we show that our framework can also take advantage of clean data to improve upon existing estimation methods.