Deepfakes Video Detection Based on Landmark Calibration and Neural Networks
Qi Fu, Zhaoda Chen, Linlin Li, Xinghe Zheng · 2024
With the rapid development of artificial intelligence technology, deep synthesis video technology has also made significant progress, which has blurred the boundaries between original videos and fake videos. Highly realistic deepfake videos cover multiple aspects such as faces, voices, and scenes, and their authenticity is difficult to discern with the naked eye. Although deep fake videos bring new fun to the entertainment field, they also pose a potential threat to personal privacy protection and social order stability. Therefore, the research on deep forgery technology and its detection technology will be a long-term topic. Currently, most deepfake video detection methods mainly focus on single-frame images and appearance features, which results in a complex model training process and is more sensitive to noise interference. How to effectively utilize the time series characteristics of videos is still an urgent problem to be solved. This paper proposes a model that combines landmark calibration and long short-term memory network (LSTM), aiming to mine the geometric and temporal features of fake videos. The model first preprocesses face images to obtain facial landmarks, and then calibrates these landmarks. The calibrated landmarks are subjected to key feature extraction through the Transformer layer, and then the LSTM layer is used to deeply mine the time series features between landmarks, thereby significantly improving the accuracy of forged video detection. This method not only improves the accuracy of detection, but also provides new perspectives and solutions for the identification of deep fake videos.