Detecting Cross-domain Deepfake Videos with Contrastive Prototype Learning

Yi Li, Plamen Parvanov Angelov · 2025

Deepfake videos are synthetic media generated using advanced deep learning techniques that manipulate or replace the visual and audio content of an original recording, enabling the creation of highly realistic yet entirely fabricated audiovisual content. The proliferation of such manipulated media poses significant societal risks, including potential misinformation, reputation damage, psychological manipulation, and erosion of trust in digital visual communication. Recent deep learning methods for deepfake detection have emerged, leveraging sophisticated machine learning models that analyze multi-modal cues, including facial inconsistencies, unnatural temporal dynamics, and visual misalignments to distinguish between authentic and synthetic content. However, these state-of-the-art detection approaches often struggle with the domain-shift challenge, where models trained on specific deepfake datasets fail to generalize effectively when confronted with unseen generation techniques or evolving synthesis technologies. To address this critical limitation, we propose a self-supervised contrastive learning framework called CPDD, introducing contrast between features and prototypes of original data to alleviate domain-specific distractions (i.e., deepfake generative models or datasets). We calculate the cosine similarity between two features or prototypes to scale the original distance, clustering the features around closely related prototypes. This process encodes the semantic structures discovered through clustering into the learned embedding space. The extensive experiments show that, compared to various benchmark deepfake detection models and domain generalization techniques, the proposed model achieves state-of-the-art performance on the cross-domain deepfake detection task across a wide range of scenarios.

Read the paper · More papers on PaperTik