Exploring LLM and Neural Networks Towards Malicious Prompt Detection

Himani S. Deshpande, Dhawal Chaudhari, Tanmay Sarode, Ayush Kamath · 2024

Atthe forefront of the Artificial Intelligence Revolution is the Generative AI domain which is making splashes in generation of new content from existing Large Language Models. Large Language Models (LLMs) are flexible, effective tools with many uses. On the other hand, they can be tricked by devious provocation. This work explores the detection of fraudulent prompts intended to produce false or damaging outputs using Long Short-Term Memory (LSTM) networks. During this study we analysed the LTSM model against a neural network model. We found a 91.69% accuracy for the LTSM model and a 92.52% accuracy for the simple neural network. However, the simple neural network had a higher F1 score of 0.73 compared to the LTSM score of 0.69. We found that the standard neural model is more efficient at identifying harmful user prompts. Our research highlights the need for taking into account various architectures of models to reduce the risk of malicious prompting in language learning models.

Read the paper · More papers on PaperTik