Graph-Based Prompt Injection Attacks Against Large Language Models
Hyeokjin Kwon, J.‐J. Kim, Wooguil Pak · 2024
As Large language models (LLMs) become widely used, injection attacks that induce violent or discriminatory responses from these models are also becoming more frequent. Initially, many attacks used elaborately designed textual prompts, but recently, with the introduction of models that support multi-modality, many attacks using images in addition to textual prompts have been attempted. Injection attacks using only existing text or images are mostly neutralized by latest safe aligned Large language models. In this paper, we present a new method of replacing sensitive words in the prompt with math expressions and performing an injection attack using a graph drawn based on them. By distributing illegal prompts across various input domains, for intrusion attacks against the latest safe aligned LLMs, our proposed approach achieves a very high success rate compared to SOTA existing injection attacks. We hope that the safety of the multimodal LLMs will be further strengthened by understanding the proposed injection attack.