RTT- or Bandwidth-Bound? Demystifying the KV Cache Transfer in Large Language Model Serving

Shengnan Yue, Mowei Wang, Yu Yan, Weiqiang Cheng, Zihan Jiang, Zhenhui Zhang · 2025

Modern large language model (LLM) serving systems increasingly adopt a prefill-decode disaggregation architecture to enhance inference efficiency. While this design improves resource utilization, it introduces latency due to the transfer of key-value (KV) cache. The community has generally assumed that this latency is bandwidth-bound and can be effectively mitigated by high-speed interconnects.

Read the paper · More papers on PaperTik