KVPR: Efficient LLM Inference with I/O-Aware KV Cache Partial Recomputation

Chaoyi Jiang, Lei Gao, Hossein Entezari Zarch, Murali Annavaram · 2025

Inference for Large Language Models (LLMs) is computationally demanding.To reduce the cost of auto-regressive decoding, Key-Value (KV) cache is used to store intermediate activations, which significantly lowers the computational overhead for token generation.However, the memory required for the KV cache grows rapidly, often exceeding the capacity of GPU memory.A cost-effective alternative is * These authors contributed equally.

Read the paper · More papers on PaperTik