DF-VLAD: Deepfake Video Detection based on Feature Aggregation

Yun Huang, Zhiming Luo, Miaohui Zhang, Wei Liu, Shaozi Li · 2021

With the rapid development of Deepfake technology, face video forgery can produce highly deceptive video content and bring serious security threats. The detection of this kind of fake video is more urgent and challenging. Most of the existing detection methods regard this problem as a common binary classification problem, and using a simple average or maximum as the prediction of video results can easily lead to missed detection or false detection. While the video-based detection work such as LSTM, in Deepfake detection, too much focus on timing modeling will affect the performance of Deepfake video detection to a certain extent. Based on this, this paper proposes a VLAD-based aggregation module DF-VLAD, which advances the aggregation of multiple frames from the output layer to the feature layer, which on the one hand makes the aggregation more flexible, on the other hand, it also uses the objective function of forgery detection to directly guide the learning of frame-level depth representation; On the other hand, this paper deals with this problem as a special fine-grained classification problem, because the difference between fake face and real face is very subtle. It is found that the existing face forgery methods such as Face2Face and NeuralTextures leave some common artifacts in the spatial domain. Different forgery methods produce different artifacts, while natural faces have more similar features. To make the model pay more attention to artifacts, a forgery trace capture model based on the fusion of self-attention mechanism and channel attention mechanism is proposed in this paper. Like other fine-grained classification methods, note intentions are used to guide the network to pay attention to key parts of the face. Experimental results on different public data sets show that the proposed method achieves the latest performance.

Read the paper · More papers on PaperTik