FakeFormer: Transformer-Based Lightweight Deepfake Video Detection Model

Xuefei Wang, Guoqiang Zhong, Qiang Song · 2025

In recent years, deepfake videos become increasingly realistic, making them nearly indistinguishable from the naked eye. The misuse of these deepfake videos poses a major threat to information security, leading researchers to develop effective deepfake video detection models. Video temporal features often contain rich and valuable information for identifying the authenticity of videos. Given that transformers have proven highly effective at modeling these features, researchers have been prompted to develop transformer-based detection models. However, due to the high-dimensional nature of video data, the efficiency of such models is often compromised, making it a challenging problem. In this paper, we propose a specialized transformer-based model for deepfake video detection, called FakeFormer. FakeFormer is a lightweight video-level detection model that maintains high efficiency, despite using a sequence of video frames as direct input. Additionally, FakeFormer fully integrates the strengths of convolutional neural networks and transformers, enabling it to effectively model both the temporal and spatial features of videos to accurately assess their authenticity. Numerous experiments validate the effectiveness of FakeFormer, including intra-dataset evaluation, cross-dataset evaluation and ablation study.

Read the paper · More papers on PaperTik