Jailbreak attack of large language model based on scene construction

Zhicong Hu, Yitong Li, Zhiyong Shen · 2024

The jailbreaking of large language models (LLMs) has recently garnered significant attention. Before releasing the model, extensive fine-tuning was conducted using RLHF (i.e., optimizing language models through reinforcement learning from human feedback) and other methods to ensure its behavior aligns with human values. However, even aligned LLMs can be maliciously manipulated, resulting in unintended behavior, termed "jailbreaking." In this study, inspired by manually crafted jailbreaking prompts, we introduce the Scenario Construction Attack (SC-Attack) and develop three seed templates for jailbreaking based on scenario construction, generating additional templates using the GPT-4o context. Additionally, to enhance the jailbreaking of aligned large models, we utilize Energy-constrained Decoding with Langevin Dynamics (COLD) to regulate the generation of adversarial suffixes for the attack. We define this combined attack as SC-COLD. Extensive experiments revealed that SC-Attack alone can achieve a 100% success rate on closed-source models like GPT-3.5 and GPT-4o, with the output content being more toxic. It also achieved a success rate of over 90% on aligned open-source models such as LLM (Llama2, Gemma). The extensive experiments demonstrate that SC-COLD possesses broad applicability, robust controllability, and a high success rate.

Read the paper · More papers on PaperTik