BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure

Large Language Models (LLMs) are increasingly deployed in modern AI infrastructure, creating a strong demand for high‐throughput and resource‐efficient serving systems. Disaggregated LLM serving, which decouples prompt prefill from auto‐regressive decode to accommodate their heterogeneous compute and memory characteristics, has emerged as a promising architecture. However, existing disaggregated serving systems suffer from three fundamental limitations: static resource allocation that fails to adapt to highly dynamic workloads, severe load imbalance between compute‐bound prefill and memory‐bound decode stages, and prefix‐cache‐aware routing that skews load distribution and creates performance hotspots. These issues collectively limit resource utilization, scalability, and the ability to meet service level objectives (SLOs) under real‐world workloads.

Paper

Similar papers

© 2026 NYSGPT2525 LLC