False Negative Masking for Debiasing in Contrastive Learning
Jen‐Tzung Chien, Kuan Chen · 2024
Contrastive learning has been popular to carry out self-supervised learning where a meaningful representation with instance discrimination is learned without any label information. However, recent studies have found that there might exist some false negative samples in training data, which have the same label as that of anchor. This phenomenon, also known as the sampling bias, considerably degrades the system performance in a downstream task. Accordingly, it is crucial to identify and reject those false negative samples without accessing their labels. This study deals with such a challenging issue and presents an approach based on the out-of-distribution (OOD) detection which identifies and masks the false negative samples. In general, the samples from the same class are seen as in-distribution (ID) data whereas the samples from the other classes are viewed as OOD data. Therefore, this study presents a new contrastive learning by detecting those false negative samples and masking them in calculation of contrastive loss during optimization. Compared with the original contrastive learning, the proposed method is illustrated with a tighter upper bound in false negative masking. The experiments demonstrate that the proposed method can achieve competitive results.