On the Dependable Operation of Key-Value Caches in Large Language Models (LLMs)
Zhen Gao, Jie Deng, Shanshan Liu, Pedro Reviriego, Fabrizio Lombardi · 2025
The use of Transformer-based architectures has triggered the fast development of large language models (LLMs); LLMs achieve unprecedented performance for a wide range of natural language processing tasks. The Attention mechanism in LLMs is computationally intensive, so most LLMs choose to cache the Keys and Values vectors of existing tokens to achieve a tradeoff between computational complexity and additional memory; this technique is generally known as KV cache. The size of this cache can be even larger than the memory storing the parameters, so, for its dependable operation as requirement in many applications, it is very important to assess its performance in the presence of soft errors in the memory. To the best of the authors' knowledge the impact of soft errors on the memory for the KV cache has not been previously studied. In this paper, the impact of bit-flip memory errors on the KV caches (with half-precision floating-point values) of two widely used LLMs (Mistral-7B and LlaMA2-7B) is evaluated based on error injection simulation. The results show that the first two exponent bits of the cache values are critical for LLM dependability, and errors on the prefilling stage tend to have a more severe impact than those on the decoding stage.