CleanerCLIP: Fine-Grained Counterfactual Semantic Augmentation for Backdoor Defense in Contrastive Learning
Yuan Xun, Siyuan Liang, Xiaojun Jia, Xinwei Liu, Jun Chen, Xiaochun Cao · IEEE Transactions on Information Forensics and Security · 2025
With the rise of the open-source community, multimodal pre-trained models such as CLIP have become increasingly vulnerable to backdoor attacks. Backdoor triggers can manipulate model outputs during inference, posing a significant threat to downstream users. While post-training defenses based on fine-tuning have made some progress, their effectiveness remains limited due to two key challenges: (1) They fail to weaken the connection between the backdoor trigger and the target text space, making it possible for the model to still rely on the backdoor pattern for prediction. (2) Although batch-level fine-tuning expands the data distribution, it lacks precise guidance for vision-language alignment. To address these limitations, we propose a fine-grained counterfactual text-driven sample-level fine-tuning defense. By generating counterfactual sub-texts, we explicitly guide the text space to shift towards a more discriminative and robust representation, thereby indirectly weakening the association between backdoor triggers and target semantics. Furthermore, we introduce intra-sample contrastive learning with hard negative sub-texts, which enforces a more precise gradient direction to enhance vision-language fine-grained alignment. We evaluate our approach against six different backdoor attack methods and conduct a comprehensive zero-shot classification study on ImageNet-1K. Experimental results demonstrate that our method surpasses SoTA CleanCLIP in defending against BadCLIP attacks, reducing the attack success rate (ASR) in Top-1 classification by 52.02% and Top-10 classification by 63.88%. We aim to enhance the robustness of multimodal models against backdoor threats, fostering safer deployment in real-world applications.