Spatiotemporal Attention-Based Deepfake Detection
Minrui Huang, Zilong Liang, Pengyang Zhang, Hengxin Li, Dongdong Zhan, Siyu Wang, MiuYi Chan · 2024
With the rapid advancement of Deepfakes technologies, generating highly realistic manipulated videos has become increasingly accessible, raising serious social and security concerns. To overcome the limitations of traditional models, such as Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs), in capturing temporal dependencies across video frames, this paper proposes a Deepfakes detection method using Spatiotemporal Attention mechanisms based on TimeSformer. Our approach detects subtle manipulations and temporal inconsistencies in videos more effectively than conventional methods. By applying transfer learning and fine-tuning in the DFDC and FaceForensics++ datasets, the proposed model achieved an AUC of 0.95 and an F1 score of 86% in FaceForensics++, along with an accuracy of 88% in DFDC. Precise facial region extraction, improved inference strategies, and optimized loss functions further improved performance and generalization across datasets. The experimental outcomes highlight the potential of Spatiotemporal Attention-Based Models for reliable video forgery detection, contributing to the authenticity and security of digital media.