MalPID: Malicious Prompt Injection Detection Dataset for Large Language Model based Applications

Sihem Omri, Manel Abdelkader, Mohamed Hamdi · 2024

Despite the significant transformation made by large language model (LLM)-based chatbots in the field of conversational artificial intelligence (AI), these systems are vulnerable to the attack known as prompt injection by malicious users to make them ignore their guardrails and generate objectionable content. However, few previous works have been addressing this issue using complex and costly mechanisms and there is a lack of large dataset that deal with prompt injection examples. In this work, we introduce MalPID, a novel dataset for malicious prompt injection detection. This benchmark contains various malicious and legitimate prompts collected from different sources and labelled manually. Our systematic evaluation of models trained on our dataset has shown impressive results in detecting prompt injection data. In the future, MalPID could be a valuable resource for advancing the creation of safe conversational AI systems.

Read the paper · More papers on PaperTik