Deepfake Video Detection Based on the Decomposition of Spatial-Temporal Attention Mechanism in ViViT

Guoqing Sun, Zhichao Lian · 2024

With the rapid development of Deepfake synthesis technology in recent years, our cybersecurity and individual privacy face challenges. In pursuit of robust Deepfake detection, researchers have tried to use temporal cues in videos, employing models such as RNNs and 3D convolutional networks. Despite these efforts, there remains ample opportunity for improvement within these models. In this paper, we introduced an approach utilizing the Video Vision Transformer, which is based on the decomposition of Spatial-Temporal attention mechanisms, for the detection of video face forgery. This method is designed to capture spatial artifacts and temporal inconsistencies. Besides, difference module is introduced to screen features and reduce the interference of natural factors on the model. To enable robust Deepfake detection. We have conducted extensive experiments across some datasets, including FaceForensics++, DFDC and WildDeepfake datasets. Which demonstrates the validity of the model.

Read the paper · More papers on PaperTik