Advancing Generalization in Deepfake Detection: Supervised Contrastive Representation Learning With Dual Stream Spatio-Temporal Features
Sang-Ho Son, Wooju Kim · IEEE Access · 2025
Deepfake technologies have rapidly evolved, enabling highly realistic facial manipulations that are increasingly difficult to detect. However, existing detection models remain limited in their ability to generalize beyond the manipulation techniques used during training. As deepfake generation methods continue to diversify, enhancing generalization has become critical for deployment in real-world scenarios. To address this challenge, this paper proposes a robust and generalizable deepfake detection framework based on supervised contrastive learning. Rather than overfitting to generation-specific artifacts, the proposed method learns discriminative representations by integrating domain- and similarity-aware contrastive loss with distributional regularization. The framework adopts a dual-stream architecture consisting of a 3D CNN (I3D with Non-Local Blocks) to capture temporal dynamics and a 2D CNN (ResNet) for spatial features. The extracted features are fused and passed to a support vector machine (SVM) classifier to refine decision boundaries. Extensive experiments on FaceForensics++, Celeb-DF, and DFDC datasets demonstrate that the proposed model achieves superior generalization performance across diverse and unseen deepfake generation techniques, outperforming existing methods in cross-dataset settings. Additionally, explainability analyses validate the model’s focus on meaningful facial regions. Appendix experiments also highlight its potential for efficient deployment with minimal performance loss.