TARDIS: A GPU-Centric KV Cache Service for Efficient LLM Inference

Yifan Hu, Shi Qiu, Jianqin Yan, Hao Chen, X. Wang, Lu Tang, Guangtao Xue, Yiming Zhang · 2025

Key-value (KV) cache is a crucial optimization for large language model (LLM) serving, particularly in long-context inference scenarios. While existing KV stores suffer from a fundamental mismatch between the CPU-centric KV cache storage and the GPU-accelerated computation: (1) the low-parallelism CPUs cannot satisfy the fine-grained, highly parallel I/O demand, and (2) the synchronization between CPUs and GPUs introduced by layerwise asynchronous KV operations prevents the adoption of low-level optimizations (such as CUDA Graph).

Read the paper · More papers on PaperTik