MM-IML: Multi-Modal Image Forgery Detection and Localization

Qing An Huang, Xiangyu Yu, Zhipei Xu · 2025

The proliferation of image editing tools like Photoshop has heightened risks of malicious image tampering, leading to societal, economic, and legal issues. Image Forgery Detection and Localization (IFDL) aims to identify tampered images and locate altered regions using mask representations. Current methods, including CNNs, ViTs, and Multimodal Large Language Models (MLLMs), face limitations such as poor generalization to unseen data, insufficient training datasets, and loss of critical details during embedding. Additionally, specialist models lack multimodal interaction and semantic-level understanding, limiting robustness. To address these challenges, we propose MM-IML, a general-specialized multimodal framework that combines the generalization capability of MLLMs with the precision of specialist models. By introducing a cross-modal tampering feature fusion module and leveraging swin transformer blocks, MM-IML enhances tampering detection and localization across modalities. Extensive experiments validate its superior performance and robustness, significantly surpassing existing methods on benchmark IFDL datasets.

Read the paper · More papers on PaperTik