A Review of the Application of Prompt Engineering in the Safety of Large Language Models

Hui Guo, Long Zhang, Xilong Feng, Qiusheng Zheng · 2024

With the rapid development of large language models, their popular application in dialogue systems has achieved certain remarkable generative performances. However, this also brings several possible security risks, such as poisoned content generation and its exploitation via adversarial prompts. Prompt engineering is one of the key technologies that guide models in generating safe and reasonable outputs by designing and optimizing prompts for both offense and defense. In this paper, LLM safety with regard to prompt engineering is systematically investigated from both attack and defense viewpoints, including common attack methods like red teaming, template-based attacks, and neural prompt attacks, and the use of system prompts, deletion/filtering mechanisms, and contrastive decoding as defense strategies. The goal is to present new perspectives on LLM safety research while predicting the possible applications and challenges facing prompt engineering in high-risk domains.

Read the paper · More papers on PaperTik