ForgeFinder: Perceptive Multimodal Deepfake Detection via Multi-grained Forgery Localization

Baoping Liu, Bo Liu, Ming Ding, Tianqing Zhu · ACM Transactions on Multimedia Computing Communications and Applications · 2025

Deepfake techniques can now generate multimodal content comprising video and audio tracks. Compared with unimodal Deepfake images, videos or audio, multimodal Deepfake content is more deceptive and easily leads to the dissemination of hate speech, incitement to violence, and disinformation. Therefore, the detection of multimodal Deepfake has attracted much research attention recently. While cross-attention shows the promising capacity for modelling the complicated dependencies between audio and video in multimodal Deepfake detection, it fails to learn accurate cross-modal patterns if audio and video are misaligned in the temporal dimension. Besides, most current multimodal Deepfake detectors only provide a binary classification label, lacking fine-grained localization to identify significant forgery in multiple dimensions (e.g., modal, time, and spatial dimension). In this study, we propose a novel multimodal Deepfake detection framework named ForgeFinder, which goes beyond binary label prediction and achieves multi-grained forgery localization in modal and spatiotemporal dimensions. ForgeFinder incorporates both intra-modal and cross-modal inconsistencies to classify multimodal input. In detail, we adopt Serial Spatiotemporal Self-Attention (SSTSA) in the Intra-Modal Inconsistency Explorer (Intra-MIE), which allows the temporal self-attention to run in the original dimension without bringing unacceptable computational complexity. In the Cross-Modal Inconsistency Explorer (Cross-MIE), we propose the Offset-Shifted Cross-Attention (OSCA) by introducing a time offset term to the conventional cross-attention to mitigate the inaccuracy of cross-modal dependencies modelling brought by the temporal misalignment. By adopting the outputs of Intra-MIE for unimodal tasks, we identify the likelihood of modals being manipulated and localize tampered modals. At the same time, the attention weights of SSTSA can be visualized to pinpoint the temporal and spatial distribution of Deepfake manipulation. Therefore, for a single audio–video input sample, ForgeFinder not only tells the authenticity of the overall input but also localizes the modal, temporal sequence, and spatial coordinates of significant forgery, significantly contributing to more comprehensive forensics analysis. The results of extensive experiments indicate that ForgeFinder achieves state-of-the-art detection performance as well as accurate forgery localization in modal and spatiotemporal dimensions. Furthermore, experiments on content generated by Diffusion Models (DMs) show that our model also effectively recognizes DM-generated content.

Read the paper · More papers on PaperTik