Automating Prompt Leakage Attacks on Large Language Models Using Agentic Approach

Tvrtko Sternak, Davor Runje, Dorian Granoša, Chi Wang · 2025

This paper introduces a novel framework for evaluating the security of large language models (LLMs) against prompt leakage-the exposure of system-level prompts or proprietary configurations-which we identify as a critical threat to secure LLM deployment. Leveraging a multi-agent system implemented using AG2 (formerly AutoGen), we design agentic teams tasked with probing and exploiting the target LLM to elicit its prompt. Inspired by cryptographic principles, we define a prompt leakage-safe system as one in which an attacker cannot distinguish between two agents: one initialized with an original prompt and the other with a prompt stripped of sensitive information. In such a system, the agents' outputs are indistinguishable, ensuring sensitive information remains secure. This framework establishes a rigorous standard for evaluating and designing secure LLMs, bridging the gap between automated threat modeling and practical LLM security through adversarial testing. The implementation of prompt leakage probing is available at GitHub: https://github.com/sternakt/prompt-leakage-probing

Read the paper · More papers on PaperTik