AFFM-FID: Attention Focus Fusion Network for Multimodal Fake Information Identification and Detection with Prompts Template
Kexin Chen, Cheng Wu Yang, Shibin Zhang · 2024
Multimodal news detection plays a critical role in identifying fake information within the vast amounts of social media data, which is essential for mitigating online risks. Although recent advancements in fake information detection have shown promise, existing methods predominantly rely on traditional feature fusion techniques that often fall short in capturing the intricate cross-modal similarities and differences between images and text. To overcome these limitations, we propose a task-oriented prompt template that leverages the textual representation capabilities of language models. Our approach introduces task-driven prompts before each tweet, which include Soft Prompts, Sentence Inputs, and Learnable Embeddings. Furthermore, we incorporate an attention focus mechanism with bidirectional attentions—both forward and backward—to enhance the capture of contextual relationships between images and text over time. Experimental results on datasets such as Twitter, Weibo, and custom collections demonstrate the effectiveness of our method, surpassing the performance of the current best model by 1.4% on Weibo and 0.4% on our Snopes dataset.