Exploring the Enhancement of Transferability of Multimodal Adversarial Examples in Vision-Language Pretraining Models
Xuan Song, Chunxian Wu, Liping Du · 2025
Vision-language pre-training models have demonstrated outstanding performance on a wide range of multimodal tasks. Nevertheless, they remain susceptible to multimodal adversarial examples. Recently, the introduction of set-level guided attack (SGA) has notably improved the transferability of adversarial samples against multimodal models. This method operates by integrating image-text pairs to enhance the variability of adversarial examples throughout the optimization procedure. Nevertheless, such a strategy may lead to overfitting on the target model, which in turn can hinder transferability. To tackle the challenge of limited transferability in visual-language adversarial attacks, this study introduces a novel attack framework. The proposed approach effectively increases the variation among adversarial instances throughout the optimization process without inducing overfitting, thereby enriching data variation. Moreover, it promotes deeper semantic integration between visual and textual modalities through graph-structured multimodal fusion guided by semantic relationships. We carried out comprehensive experiments using publicly available datasets. The results demonstrate that, across all evaluated tasks, our method attains significantly superior transferability when compared to SGA and other benchmark approaches.