"Prompter Says": A Linguistic Approach to Understanding and Detecting Jailbreak Attacks Against Large-Language Models

Dylan Lee, Shaoyuan Xie, Shagoto Rahman, Kenneth Pat, David Lee, Qi Alfred Chen · 2023

Large language models (LLMs) designed for safety and harmlessness remain vulnerable to adversarial exploitation. This susceptibility is evidenced by the frequent occurrence of "jailbreak'' attacks, which successfully induce undesired behaviors using carefully designed prompts. This study investigates how to distinguish between safe and harmful prompts for LLMs using linguistic analysis. We first assemble a comprehensive dataset of labeled prompts (benign vs. malicious) from existing research. By analyzing the syntactic, lexical, and semantic features of these prompts, we developed a rubric to identify prompt intent solely from text. This rubric inspired the creation of a machine learning model that incorporates these features. We tested various machine learning algorithms, including logistic regression, support vector machines, and multi-layer perceptrons to understand how different feature representations interact with each model type. Our results reveal optimal combinations of classifiers and features for preemptively flagging malicious prompts before they reach LLMs. This model serves as a foundational tool, embedding the principles of our rubric for automated malicious prompt detection. Furthermore, we explore the linguistic differences across various languages, examining semantic propensity, textual structure, and syntactic variations, and how these impact the effectiveness of jailbreaking attempts on multiple LLMs. This research contributes to the significant problem of LLM security and reliability, paving the way for future advancements.

Read the paper · More papers on PaperTik