Defending Against GCG Jailbreak Attacks with Syntax Trees and Perplexity in LLMs
Qiuyu Chen, Shingo Yamaguchi, Yudai Yamamoto · 2024
In this paper, we propose a novel classification method that utilizes syntax trees and perplexity to identify jailbreak attacks that use hostile suffixes to make large language models (LLMs) more likely to generate dangerous content, such as Greedy Coordinate Gradient (GCG) attacks. GCG jailbreak attacks deceive the model with hostile suffixes, prompting LLMs to produce responses that exceed ethical limits. To address this issue, we propose a classifier that can confirm whether the input to the LLM contains a suffix indicative of a GCG jailbreak attack. The STPC (Syntax Trees and Perplexity Classifier) introduces syntax trees to analyze the structural characteristics of GCG suffixes and combines perplexity for comprehensive classification judgment. It accurately detects most jailbreak attacks in the test set, mitigates false positives, and improves the reliability of LLM responses. In all STPC test results, the classification accuracy reached 96.2%.