LLM-Sentry: A Model-Agnostic Human-in-the-Loop Framework for Securing Large Language Models
Saquib Irtiza, Khandakar Ashrafi Akbar, Arowa Yasmeen, Latifur Khan, Ovidiu Daescu, Bhavani M. Thuraisingham · 2024
LLM-Sentry represents a novel black-box defense strategy to safeguard Large Language Models (LLMs) against jailbreaking attacks. A key advantage of our approach is its model-agnostic nature, as it does not rely on specific information about the model’s architecture or parameters, thereby enabling its application to any commercial or open-source language models. Additionally, LLM-Sentry does not require retraining when new jailbreak attacks are discovered; a simple update to the knowledge base equips LLM-Sentry to defend against new threats. The widespread adoption of LLMs is attributed to their high-quality responses and user-friendly nature. However, these models are susceptible to manipulation by malicious actors exploiting vulnerabilities to generate harmful or compromised content. Recent research has identified various jailbreaking methods that exploit vulnerabilities in LLM security measures. Given the increasing complexity of jailbreaking techniques and the ambiguous nature of LLM safeguards, it is imperative to develop unique defense strategies that can seamlessly integrate into existing security frameworks and make them robust.Our work comprehensively analyzes various commercial LLMs to assess their vulnerability to sophisticated, multilingual jailbreaking prompts. We propose a defensive approach that combines a Zero-shot language classifier with the Retrieval Augmented Generation (RAG) technique to screen and filter potentially harmful input prompts before they are processed by the language model for response generation. We adopt a human-in-the-loop approach to gather a dataset comprising harmful and safe prompts, which serves as a knowledge base for the RAG retriever module to extract relevant context. Our investigation includes successful jailbreaking attempts on prominent commercial LLMs like Gemini, Mistral 7B, and ChatGPT, wherein we successfully bypass existing security measures and elicit compromised responses. We conduct a thorough evaluation of our approach against various baseline methods to validate its resilience and superiority against such attacks empirically. Our approach achieves an attack detection accuracy of 97%, surpassing all other methods in our comparative analysis.