VRLVMix: Combating Noisy Labels with Sample Selection based on Loss Variation

Hailun Wang, Zhengzheng Tu, Bo Jiang, Yuhe Ding · 2023

Since deep neural networks can fully fit all data, including noisy labels, that is, mislabeled data, this will be detrimental to the robustness and generalization ability of the network. To address this problem, existing methods usually use small loss tricks to select clean samples for training. However, sample selection using only small loss techniques cannot distinguish between large loss “hard” samples and noise samples. In this work, we analyze the reasons why clean samples are misidentified as noise samples and propose a balanced selection mechanism based on loss variation. This method sorts the amount of loss change, uses variance to count the rankings of multiple epochs to describe the stability of loss variation, and selects samples with stable loss as clean samples to reduce the misclassification of “hard” samples. At the same time, semantic clustering is used to assist resampling and reweighting, thereby alleviating the negative impact1of class imbalance on sample selection. We conduct comparative experiments and ablation experiments on synthetic noise datasets and real-world datasets such as CIFAR-10/100 and Clothing1M. Our VRLVMix has been shown to outperform numerous state-of-the-art methods, as evidenced by the experimental results. Moreover, it demonstrates the ability to extract clean samples even from large loss sample.

Read the paper · More papers on PaperTik