Bypassing Guardrails: Lessons Learned from Red Teaming ChatGPT

Terry Yue Zhuo, Yujin Huang, Chunyang Chen, Xiaoning Du, Zhenchang Xing · ACM Transactions on Software Engineering and Methodology · 2025

Ethical and social risks persist as a crucial yet challenging topic in human-AI interactions, especially in ensuring the safe usage of natural language processing (NLP). The emergence of large language models (LLMs) like ChatGPT introduces the potential for exacerbating this concern. However, prior works on the ethics and risks of emergent LLMs either overlook the practical implications in real-world scenarios, lag behind rapid NLP advancements, lack user consensus on ethical risks, or fail to holistically address the entire spectrum of ethical considerations. In this article, we comprehensively evaluate, qualitatively explore, and catalog ethical dilemmas and risks in ChatGPT through benchmarking with eight representative datasets and red teaming involving diverse case studies. Our findings show that while ChatGPT demonstrates superior safety performance on benchmark datasets, its guardrails can be bypassed via our manually curated examples, revealing not only the limitations of current benchmarks for risk assessment but also unexplored risks in five distinct scenarios, including social bias in code generation, bias in cross-lingual question answering, toxic language in personalized dialogue, misleading information from hallucination, and prompt injections for unethical behaviors. We conclude with implications from red teaming ChatGPT and recommendations for designing future responsible large language models.

Read the paper · More papers on PaperTik