Attention-Guided Supervised Contrastive Learning for Deepfake Detection
Saima Waseem, Syed Abdul Rahman bin Syed Abu Bakar, Bilal Ashfaq Ahmed · 2024
Recent advancements in face deepfake detection have shown impressive results. However, prior studies typically used crossentropy loss to approach face manipulation detection as a classification problem. Approaches based on cross-entropy loss prioritize category distinctions over capturing the underlying differences between real and fake faces, which restricts the model's capacity to generalize to unseen datasets. As original image or video can closely resemble the deepfake in terms of appearance, making it challenging to distinguish them, we propose to utilize the differences in the representation space to develop a generalizable detector. In this paper, we present an attention-guided supervised contrastive learning approach for deepfake detection, aiming to leverage differences in the representation space and prioritize disparities between classes rather than focusing solely on categories. By using supervised contrastive learning, the model learns to create a discriminative representation by contrasting between classes, while an attention module directs the model to relevant features for each class and filters out irrelevant features. This method learns features from a wide variety of deepfake images, thus improving the accuracy of deepfake detection in unseen datasets. Experimental results show the effectiveness of our attention-guided supervised contrastive learning deepfake detector on benchmark datasets such as FF++, Celeb-DF, DFD and DFDC-P.