Enhancing System Security: LLM-Driven Defense Against Prompt Injection Vulnerabilities
Oleksandr Muliarevych · 2024
This article examines cybersecurity vulnerabilities in systems utilizing Language Model Interfaces, focusing on the challenges of building secure systems. It provides an overview of current interfaces and their associated risks. A key contribution is the design of a prompt analysis and injection detection subsystem, which assesses input relevance and security. The integration of an additional Filter level in the system safeguards against prompt injection attacks by pre-processing user requests and post-processing responses. This level includes a Prompt Analyzer for input wrapping, and an Attack Validator module using the LLM to classify prompts. The study evaluates various prompt injection attacks across different system setups, including core GPT -3 and GPT-4 models without additional security layers, a Fuzzy Search method, and systems using an Attack Validator based on GPT-3 and GPT -4 models.