Optimizing Transformer Models for Prompt Jailbreak Attack Detection in AI Assistant Systems

Le Tien, Phạm Văn Hưởng · 2024

Integrating AI-powered assistants in different areas has changed how people retrieve information. Users can now get quick responses, use the service whenever needed, and handle inquiries well. However, relying on these systems has increasingly raised worries about their safety and reliability, especially when facing threats like prompt injection or jailbreak attacks. These attacks take advantage of the weaknesses of large language models (LLMs) by changing input prompts to create harmful, biased, or misleading results. This paper examines prompt jailbreak attacks in ChatGPT-based assistant chatbots, especially in education. It examines the weaknesses in these systems and suggests ways to detect these attacks by fine-tuning various Transformer models. The research also includes adding prompt jailbreak attack detection in a virtual assistant application for university use. This aims to ensure teachers and students can interact safely and rely on the system. Through rigorous experimentation and evaluation, we demonstrate the practicality and effectiveness of our detection methods, providing reassurance and confidence in the face of AI security challenges. Our studies emphasize the need for robust AI chatbots and offer practical solutions to preserve the integrity of AI-driven tools.

Read the paper · More papers on PaperTik