Frequency-Assisted Temporal Upsampling Artifacts Representation Learning for Face Forgery Detection

Rui Sun, Xiaolu Yu, Fei Wang, Zikai Da, Yifan Zhang, Jun Gao · IEEE Transactions on Biometrics Behavior and Identity Science · 2025

Recently, the highly realistic deepfake images have aroused the potential for deepfake abuse, and both the detection and its generalization remain challenging. Existing deepfake video detection methods strive to capture distinguishing features between fake and real faces through temporal modeling. However, these works provide local supervision between sampled video frames, but overlook intra-frame generative artifacts, which can function as an useful indicators for detection. To mitigate the issue, we propose a temporal-frequential convolutional network based on the representation of upsampling artifacts, which can automatically capture incoherent artifacts representation in different ends. Specifically, we design the Upsampling Artifacts Representation Module (UARM), which captures upsampling artifacts introduced in generation end based on the relationship between pixels, and then we probe the deep coupling of intra-frame global frequency information with inter-frame incoherence, and design the Frequency-Assisted Temporal Incoherence Module (FATIM), which can facilitate the detection of temporal information in RGB domain by frequency domain information, and the Self-tuned Convolutions Module (SCM) is designed to supplement information and adjust it automatically in the learning process. In addition, the Information Fusion Module (IFM) is designed to couple multiple information and establish a more comprehensive representation. With the above modules, our proposed method can learn robust features in both generation end and detection end. Extensive experiments demonstrate that our method outperforms other state-of-the-art video detection methods by a large margin on the challenging DFDC and FF++ in low-quality.

Read the paper · More papers on PaperTik