FST-Net: Exploiting Frequency Spatial Temporal Information for Low-Quality Fake Video Detection
Min Zhang, Xiaohan Liu, Chenyu Liu, Xueqi Zhang, Haiyong Xie · 2021 IEEE 33rd International Conference on Tools with Artificial Intelligence (ICTAI) · 2021
Recently, state-of-the-art face manipulation algorithms have made significant progresses to forge images and videos that are able to deceive human eyes or even detection algorithms, which brings new challenges to forgery detection. In particular, the performance of detection algorithms for forgery videos is not as perfect as that for forgery images; furthermore, the difficulty of detection increases dramatically for low-quality videos. To address this challenge, we propose a novel dual stream architecture, referred to as FST-Net, for jointly mining forged features in the frequency, spatial and temporal domains. Specifically, we extract the spectral information of different frequency bands to expose intra-frame artifacts, and use the separable 3D CNN (S3D) to extract the spatio-temporal features among video frame groups. Moreover, to make the model focus on the tampered area, we add an attention layer to both backbone networks. Comprehensive experiments show that our model outperforms existing methods in video detection on challenging FaceForensics++ datasets, especially on low-quality video datasets.