Feedback-Based Modal Mutual Search for Enhanced Adversarial Attacks on Vision-Language Pre-Training Models
Renhua Ding, Xinze Zhang, Xiao Bo Yang, Kun He · IEEE Transactions on Emerging Topics in Computational Intelligence · 2026
Despite significant advancements in vision-language pre-training (VLP) models for cross-modal tasks, their vulnerability to adversarial attacks remains a critical issue. Among the prevailing attack strategies, transfer-based black-box attacks, which leverage data augmentation and cross-modal interactions to generate transferable adversarial examples on surrogate models, have shown promise in real-world scenarios. However, these methods often face challenges in attack transferability, stemming from discrepancies in feature representations across different models. In this paper, we propose a novel attack paradigm, called feedback-based modal mutual search (FMMS), to address this limitation. FMMS introduces a modal mutual loss (MML) that encourages the feature space to push matched image-text pairs apart while simultaneously pulling mismatched pairs closer, driving the adversarial examples toward more effective perturbations. Furthermore, FMMS leverages feedback from the target model, enabling iterative refinement of adversarial examples and progressively steering them into the adversarial region. This feedback mechanism effectively bridges the feature representation gap between surrogate and target models, resulting in more potent and transferable adversarial examples. Extensive evaluations on the Flickr30K and MSCOCO datasets for image-text matching tasks demonstrate that FMMS significantly outperforms existing state-of-the-art baselines. Additionally, combining FMMS with data augmentation further enhances the attack performance, highlighting its substantial potential in undermining VLP models in practical settings.