Diffusion-Based Adversarial Generation with SAM-Guided Spatial Semantics for Text-to-Image Models

Zhanghao Qin · 2025

While adversarial examples have been extensively studied, most mainstream approaches focus on perturbations in the pixel space, which often introduce high-frequency artifacts and suffer from limited cross-model transferability. Moreover, these pixel-level modifications typically ignore the semantic structure of the image, reducing the effectiveness of targeted attacks. To address these limitations, we explore adversarial generation in the latent space of diffusion models, guided by semantic masks from the Segment Anything Model (SAM). In this work, we propose an innovative strategy that leverages large-scale text-to-image generative models. By integrating adversarial optimization with reversible diffusion inference (DDIM Inversion), we produce adversarial examples that exhibit both high fidelity and robust attacking performance. Meanwhile, we employ the Segment Anything Model (SAM) for semantic parsing to deliver fine-grained spatial guidance during the text-to-image generation process, enabling more effective perturbations in local regions.Extensive experiments demonstrate that our proposed method significantly enhances the adversarial potential of generated examples across diverse architectures and defense settings. For instance, on black-box models such as Swin-B and DeiT-B, our method achieves up to 45.8% and 51.4% attack success rates, respectively, outperforming other diffusion-based methods like DiffAttack by margins of 5% and 1.5%, while maintaining superior perceptual fidelity. These results confirm the effectiveness of semantic-aware latent perturbations in boosting cross-model transferability and imperceptibility.

Read the paper · More papers on PaperTik