STDD: A Hybrid Spatial-Temporal model for Deepfake Detection

Thuan Minh Pham, Lam Thu Bui, Hien Van Tran, Trung Duy Pham · 2025

The rise of deepfake videos has posed significant challenges to digital media verification, requiring robust methods for detecting both spatial and temporal inconsistencies. In this paper, we propose STDD, a hybrid deepfake detection model that combines Convolutional Neural Networks (CNNs) and Transformers to analyze both spatial and temporal features in video frames. Specifically, CNNs are used for initial spatial feature extraction, while both self-attention mechanisms of the Transformer capture spatial dependencies, and cross-attention mechanisms handle temporal relationships across frames. We evaluate STDD on the DFDC dataset, demonstrating that our model outperforms existing methods in terms of accuracy and robustness. The results suggest that the integration of spatial and temporal analysis through this hybrid approach leads to significant improvements in deepfake detection.

Read the paper · More papers on PaperTik