BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure

Yan‐Lin He, Minxian Xu, Jingfeng Wu, Jianmin Hu, Chong Ma, Min Xuan Shen, Le Chen, Chengzhong Xu, Lin Qu, YE Ke-jiang · Software Practice and Experience · 2026

ABSTRACT Objective Large Language Models (LLMs) are increasingly deployed in modern AI infrastructure, creating a strong demand for high‐throughput and resource‐efficient serving systems. Disaggregated LLM serving, which decouples prompt prefill from auto‐regressive decode to accommodate their heterogeneous compute and memory characteristics, has emerged as a promising architecture. However, existing disaggregated serving systems suffer from three fundamental limitations: static resource allocation that fails to adapt to highly dynamic workloads, severe load imbalance between compute‐bound prefill and memory‐bound decode stages, and prefix‐cache‐aware routing that skews load distribution and creates performance hotspots. These issues collectively limit resource utilization, scalability, and the ability to meet service level objectives (SLOs) under real‐world workloads. Methods To address these challenges, we propose BanaServe, a dynamic orchestration framework for disaggregated LLM serving that continuously rebalances both computational and memory resources across prefill and decode instances. BanaServe introduces three key mechanisms: (i) layer‐level weight migration to enable coarse‐grained redistribution of computation, (ii) attention‐level Key–Value (KV) cache migration for fine‐grained memory load balancing, and (iii) a Global KV Cache Store with layer‐wise overlapped transmission to decouple routing decisions from cache placement. Together, these mechanisms eliminate cache‐induced hotspots and allow routers to perform purely load‐aware scheduling with minimal latency overhead. BanaServe is implemented on top of state‐of‐the‐art LLM serving frameworks, including vLLM and DistServe. Results We evaluate BanaServe under diverse and challenging workloads, including long‐context inference, bursty request arrivals, and mixed prompt–generation patterns. Experimental results show that, compared to vLLM, BanaServe improves throughput by 1.2–3.9× and reduces total processing time by 3.9%–78.4%. In comparison with DistServe, BanaServe achieves 1.1–2.8× higher throughput while reducing latency by 1.4%–70.1%. These gains are consistent across workload variations, demonstrating BanaServe's robustness under highly dynamic serving conditions. Conclusion BanaServe demonstrates that dynamic, multi‐granularity resource rebalancing and cache‐decoupled routing are essential for efficient disaggregated LLM serving. By jointly addressing resource elasticity, stage imbalance, and cache‐induced load skew, BanaServe substantially improves throughput, latency, and resource utilization in real‐world deployments. This work provides a practical and scalable foundation for next‐generation LLM serving systems operating under dynamic and heterogeneous workloads.

Read the paper · More papers on PaperTik