Medical VLP Model Is Vulnerable: Toward Multimodal Adversarial Attack on Large Medical Vision-Language Models
Zimu Lu, Ning Xu, Hongshuo Tian, Lanjun Wang, An-An Liu · IEEE Transactions on Circuits and Systems for Video Technology · 2025
Medical Visual Question Answering (Medical VQA) is an essential task that facilitates the automated interpretation of complex clinical imagery with corresponding textual questions, thereby supporting both clinicians and patients in making informed medical decisions. With the rapid progress of Vision-Language Pretraining (VLP) in general domains, the development of medical VLP models has emerged as a rapidly growing interdisciplinary area at the intersection of artificial intelligence (AI) and healthcare. However, few works have been proposed to evaluate the adversarial robustness of medical VLP models, which faces two primary challenges: (1) the complexity of medical texts, stemming from the presence of terminologies, poses significant challenges for models in comprehending the text for adversarial attack; (2) the diversity of medical images arises from the variety of anatomical regions depicted, which requires models to determine critical anatomical regions for attack. In this paper, we propose a novel multimodal adversarial attack generator for evaluating the robustness of medical VLP models. Specifically, for the complexity of medical texts, we integrate medical knowledge when crafting text adversarial samples, which can facilitate the terminologies understanding and adversarial strength; for the diversity of medical images, we divide the anatomical regions into either global or local regions in medical images, which are determined by learned balance weights for perturbations. Our experimental study not only provides a quantitative understanding in medical VLP models, but also underscores the critical need for thorough safety evaluations before implementing them in real-world medical applications.