Finding label noise examples in large scale datasets
Rajmadhan Ekambaram, Dmitry B. Goldgof, Lawrence Hall · 2017
Mislabeled examples are difficult to avoid while building large scale datasets. In this paper we discuss an efficient approach for finding those mislabeled examples. Our approach involves selecting a small number of potentially mislabeled examples for review by an expert. We demonstrate the utility of our method by finding some mislabeled examples in one large scale dataset (ImageNet). We found 92 errors by automatically selecting 3607 examples to review out of 22951 images from 18 classes. This requires reviewing 9 times fewer examples than the random sampling method to find an equivalent number of mislabels.