Defending Large Language Models Against Jailbreak Attacks Through Chain of Thought Prompting

Yanfei Cao, Naijie Gu, Xinyue Shen, Daiyuan Yang, Xingmin Zhang · 2024

With the deep research and widespread application of Large Language Models (LLMs), the security and privacy issues inherent in them have gradually become prominent, posing new challenges in the field of network security. Elaborately designed jailbreak attack prompts may induce LLMs to act against human values and preferences and generate harmful responses. To address this issue, we introduce a straightforward yet effective defence mechanism: Chain of Thought Prompting. This defence mechanism aims to emulate human thought processess. Chain of Thought Prompting method fully harnesses the inherent reasoning ability of the LLMs via five stages. It encourages Large Language Models (LLMs) to engage in self-thought, self-reflection, and self-refinement. Our experiments on ChatGLM-3-6B and Llama-2-7B-Chat, utilizing the JADE and DAN datasets, demonstrate a significant reduction in the average Attack Success Rate (ASR) for jailbreak attacks. This reduction is from 65.01% to 13.40% for ChatGLM-3-6B, and from 49.06% to 0.13% for Llama-2-7B-Chat. Our findings suggest that Chain of Thought Prompting can significantly reduce the generation of harmful content. Our work provides an effective method for LLMs to counter jailbreak attacks, enhancing the safety of LLMs without the need for additional training.

Read the paper · More papers on PaperTik