DeepStream-X: A Two-Stream Deepfake Detection Framework using Spatiotemporal and Frequency Features

Nada Mohamed Emara, Mazen Nabil Elagamy · 2024

Deepfake videos are increasingly becoming a major concern in digital media, driven by the rapid advancements in DL-based generation techniques. These synthetic videos, which often feature highly realistic alterations of a person’s face, pose significant threats in areas such as disinformation, political manipulation, and personal defamation. Early deepfake detection methods primarily relied on identifying visual artifacts in video frames. The majority of deepfake datasets introduced visible spatial artifacts, making these methods quite effective. However, with the advancement of deepfake generation methods, visual artifacts are diminishing, challenging spatialbased methods. As a result, recent research has shifted towards utilizing spatiotemporal, frequency, or correlation features that capture invisible artifacts and subtle inconsistencies caused by generation methods. However, many of those methods fail to generalize to more realistic datasets such as Celeb-DF. Consequently, this research proposes DeepStream-X, which is a two-stream deepfake video detection framework that fuses frequency and spatiotemporal features. The proposed framework examines the intra-frame subtle changes through frequency analysis and inter-frame motion inconsistencies through dense optical flow motion estimation. The fused features are passed to a modified Xception with attention to enhance performance. DeepStream-X is evaluated on Celeb-DF, which is a challenging realistic high-quality dataset. The proposed framework yielded an accuracy of 95.24% and an AUC of 99.45% outperforming existing state-of-the-art methods.

Read the paper · More papers on PaperTik