SCALPEL: Bypassing LLM Safety Through the Unguarded KV Cache
Tianyu Lu · Zenodo (CERN European Organization for Nuclear Research) · 2026
The KV cache—the dominant stateful component of LLM inference—is an architectural blind spot: safety-critical signals propagate from aligned weights into this mutable runtime buffer, yet no production framework verifies the integrity of cached value vectors. We exploit this gap with SCALPEL, an attack that identifies safety-critical directions in the model's internal representations, then surgically erases them from cached values at every decoding step. The attack requires no adversarial prompts, no weight modification, and no fine-tuning; it is ephemeral (active only for a single request) and invisible under the deployed integrity checks we evaluated—weight hashing, forward-pass hook auditing, and behavioral spot-checks all pass under active attack. Across eight models (7B–70B) from five vendors, SCALPEL achieves 59–81% classifier-reported attack success rate on 7–9B models with minimal calibration (33 contrastive pairs) under all system prompt conditions, rising to 77–86% with optimized directions (vs. 80–92% for weight abliteration on Llama-3-8B) without modifying any model parameters. Against Circuit Breakers, the leading representation-level defense, even best-case weight abliteration achieves 0% while cache erasure reaches 49.1%, demonstrating a purely mechanism-driven gap. On vLLM—the dominant production serving framework—approximately 30 lines of framework-specific code achieve equivalent results with 86% cross-framework agreement. A security audit of five major inference frameworks confirms that none verify the integrity of cached value vectors, establishing cache-level safety bypass as a cross-framework integrity gap.