Disrupting Large Language Models with Hidden Prompt Injection Attacks Embedded in HTML Pages

Ionuţ-Vlăduţ Dinu, Gabriel Mihail Danciu, Raul Cristian Vintilă, Titus Bălan · 2025

This research evaluates the reliability of Large Language Models (LLMs) as collaborative tools for extracting information from various web sources, including standard websites, e-commerce platforms, blogs, landing pages, or even e-mails, where the content is much shorter and the processing language models have a higher chance of getting attacked. Prior to conducting model assessments, we present some methodologies involved in web-scraping approaches, prompt injection attack types, and datasets that evaluate the attack success rate (ASR) of LLMs. Our investigation subjects LLMs to rigorous testing using HTML pages containing concealed instructions designed to terminate response generation when users make queries. The evaluation dataset consists of HTML documents of varying sizes, containing randomly positioned prompt injection attacks programmed to direct the LLM to respond with a single specific key word. Using this benchmark, we assess locally deployed LLM models, revealing their ASR across multiple architectures. The defensive mechanism we propose demonstrates a significant reduction in ASR through the implementation of a filtering system with fixed-sized windows of tokens that analyzes HTML content and eliminates compromised data, reducing the success rate of the attacks, which proved to reduce the ASR.

Read the paper · More papers on PaperTik