Paper Summary Attack: Jailbreaking LLMs Through LLM Safety Papers
Liang Lin, Songlin Hu, Xuehai Tang · 2026
The safety of large language models (LLMs) has garnered significant research attention. In this paper, we argue that previous empirical studies demonstrate LLMs exhibit a propensity to trust information from authoritative sources, such as academic papers, implying new possible vulnerabilities. Based on this insight, a novel jailbreaking method, Paper Summary Attack (PSA), is proposed. It systematically synthesizes content from either attack-focused or defense-focused LLM safety papers to construct an adversarial prompt template, while strategically infilling harmful queries as adversarial payloads within predefined subsections. Extensive experiments show significant vulnerabilities not only in base LLMs, but also in state-of-the-art reasoning models like Deepseek-R1. Our findings highlight the urgent need to reassess LLMs’ trust mechanisms in processing authoritative sources.