Graph-Based Filtering to Prevent Prompt-Engineered LLM Training Data Leaks

Alan Barnett, Seán Ahearne, Paul Barry, Merry Globin, Colin Duggan · 2025

Machine-learning generative Artificial Intelligence tools, specifically large-language models, provide varied functionality, like content generation, user-facing chatbots, and code generation. The LLM typically works with a decision engine, such as a neural network. LLMs suffer issues with training data poisoning, copyright of generated content, and this paper's focus; prompt engineering attacks and training data leaks. The authors propose an architecture to co-locate a filtering mechanism with the LLM chatbot to identify and preventing disclosure of leaked LLM training data before communication to the end-user. Implementation of a resource description framework (RDF) based filtering mechanism compares LLM outputs against a bank of training data using three approaches; the first uses a bank of hash-codes generated from training data artifacts, the second uses a bank of training data stored as plaintext, and the third couples natural language processing (NLP) with the plaintext training data bank. Accuracy, overhead and acceleration results are detailed, and observed anomalies in LLM responses to testing including plausible leaks are also discussed.

Read the paper · More papers on PaperTik