Deepfake Video Detection Using Facial Feature Points and Ch-Transformer
Rui Yang, Rushi Lan, Zhenrong Deng, Xiaonan Luo, Xiyan Sun · ACM Transactions on Multimedia Computing Communications and Applications · 2024
With the development of Metaverse technology, the avatar in Metaverse has faced serious security and privacy concerns. Analyzing facial features to distinguish between genuine and manipulated facial videos holds significant research importance for ensuring the authenticity of characters in the virtual world and for mitigating discrimination as well as preventing malicious use of facial data. To address this issue, the Facial Feature Points and Class-head-Transformer (FFP-ChT) deepfake video detection model is designed based on the clues of different FFPs distribution in real and fake videos and different displacement distances of real and fake FFPs between frames. The face video input is first detected by the BlazeFace model, and the face detection results are fed into the FaceMesh model to extract 468 FFPs. Then, the Lucas–Kanade (LK) optical flow method is used to track the points of the face, the face calibration algorithm is introduced to re-calibrate the FFPs, and the jitter displacement is calculated by tracking the FFPs between frames. Finally, the Ch is designed in the transformer, and the FFPs and FFP displacement are jointly classified through the ChT model. In this way, the designed ChT classifier is able to accurately and effectively identify deepfake videos. Experiments on open datasets clearly demonstrate the effectiveness and generalization capabilities of our approach.