Broken Bags: Disrupting Service Through the Contamination of Large Language Models With Misinformation
Yonghua Mo, Maoyang Tang, Ruohan Lin, Bohao Zhou, Xiaojian Li · IEEE Access · 2025
Large language models (LLMs) have progressively become essential production tools in contemporary society, owing to their formidable natural language generation and contextual reasoning skills. To facilitate the development of current responses by LLMs, individuals have used retrieval-augmented generation (RAG) technology, which extracts material from the corpus to assist large language models in producing relevant replies. The extensive utilization of huge language models necessitates urgent RAG security research. Conventional RAG attack techniques exhibit inadequate hiding and a substantial volume of harmful messages. Consequently, we have introduced an innovative attack mechanism termed "Broken Bags," which adeptly injects a minimal quantity of toxic text to mislead large language models. The attack is executed through a hybrid approach that incorporates artificial prompt templates, toxic content generated by LLMs, and filtering mechanisms. For instance, when the RAG system engages with publicly available knowledge bases, adversaries can take advantage of the accessibility of these RAG knowledge bases to introduce malicious texts into the retrieval database, so as to intentionally alter the model’s behavior. This work employs the linguistic similarity between toxic content and the geographical vector characteristics of the "query question" to influence the information returned by RAG, hence preventing the LLM from generating responses to the target questions. We developed and refined an artificial prompt template to render toxic language more akin to authentic human expressions and less detectable. Experimental data indicates that our attack success rate attains 94%. Ultimately, we systematically evaluate state-of-the-art defenses (including perplexity-based detection and knowledge extension, among others), and the findings indicate that these measures are unable to counter "Broken Bags," hence significantly enhancing the success rate of assaults on RAG systems.