Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques

Neusha Javidnia, Bita Darvish Rouhani, Farinaz Koushanfar · 2025

Large language models (LLMs) have demonstrated exceptional ca-pabilities in generating text, images, and video content. However, as context length grows, the computational cost of attention increases quadratically with the number of tokens, presenting significant effi-ciency challenges. This paper presents an analysis of various Key-Value (KV) cache compression strategies, offering a comprehensive taxonomy that categorizes these methods by their underlying princi-ples and implementation techniques. Furthermore, we evaluate their impact on performance and inference latency, providing critical in-sights into their effectiveness. Our findings highlight the trade-offs involved in KV cache compression and its influence on handling long-context scenarios, paving the way for more efficient LLM implemen-tations.

Read the paper · More papers on PaperTik