AIAI EngineerAug 27, 2026· 21:48

KV Cache-Aware Routing and P/D Disaggregation on Kubernetes — Yuchen Fama & Ashish Kamra, Red Hat

Yuchen Fama and Ashish Kamra of Red Hat argue that public inference benchmarks hide the chaotic reality of agentic workloads, and show how KV cache-aware routing plus prefill/decode disaggregation in the open source llm-d framework tackles it. Red Hat's traces show agentic sessions running from a few turns to 3,000, cache hit rates over 90%, and input-output token ratios past 100:1, making a 10x cost gap between cached and uncached tokens. A live demo shows routing reusing cache on the same pod cutting time from 3 seconds to 1, while a fresh system prompt pays the full 3 again. For prefill/decode disaggregation, they report P99 inter-token latency dropping from roughly 900 milliseconds to about 100 across 16 H100s serving gpt-oss, but note it only wins in the middle concurrency band and requires RDMA or RoCE to move KV caches. Their closing case study runs GLM 5.2 on H200s with three prefill workers to one decode, achieving 4x faster time to first token and 60% more requests.

  1. 0:00Intro
  2. 3:25Agentic traces
  3. 5:02Cache economics
  4. 6:24Cache routing
  5. 7:42Demo
  6. 9:29Disaggregation
  7. 12:50PD results
  8. 15:37PD tradeoffs
  9. 18:12Case study
  10. 20:26Future

Powered by PodHood

Transcript

Intro0:00

Ashish Kamra0:13

Alright, um, welcome everyone to yet another inference talk. I hope you have had a good conference so far, and, uh, so in this session—I mean, I'm sure you people who have been in the room must have heard these terms many times by now—so we're going to do a little bit more deep dive into the challenges of LLM deployments for agentic workloads.

And, uh, in this session we'll focus specifically on KV Cache-Aware Routing and, uh, P/D Disaggregation. Um, and also, you know, when you—when you look at public inference benchmark results, you are typically looking at very steady-state, isolated, highly sanitized numbers.

And what those benchmarks actually don't show you, uh, is the chaotic reality of multi-turn interactions, massive context fluctuations, which are very typical of agentic workloads. So we'll also, uh, try to pull the curtain back on some of those complexities.

Um, by way of introduction, my name is Ashish Kamra. I'm a senior manager of performance engineering at Red Hat. And with me—

Yuchen Fama1:19

Hi, I'm Yuchen. I'm the product manager at Red Hat inference, working closely with, uh, VOM and LMD core maintainers, also a contributor myself.

Ashish Kamra1:29

So here is the agenda for the next 20 minutes or so. Um, Yuchen will start with an analysis of inference behavior in the agentic era and some of the core characteristics and challenges. Uh, next we—next you will walk us through the KV Cache, um, utilization and management strategies.

I will break down the mechanics of prefill, decode, disaggregation, and walk you through some, some results. And then Yuchen will, again, bring it all back together with our ongoing case study on our favorite open coding model, GLM 5.2.

Um, and just a couple of, uh, sources from our side. If you are more interested in learning more about open source inference, we have a free course, free course on deep learning.ai, uh, by Cedric and with Andrew Ng.

Um, and the other is a series of blogs on the Red Hat developer portal on distributed inference concepts, uh, troubleshooting, and deployment patterns.

Uh, and for those who may not be aware, since Red Hat is better known as the Linux company for enterprise Linux and, uh, the Kubernetes company for OpenShift, uh, but more recently we are also a major player in open source AI inference, with, uh, us being the top contributor in VLLM, LLMD, and the CASE of projects, and also, uh, having incubated Guide LLM for benchmarking, LLM compressor for model quantization, and speculators for, uh, speculative, uh, decoding models.

And we also bring it—bring all of that together in an optimized model hub on Hugging Face under the redhat.ai.org. Uh, and we are also building the platform for the next wave of agentic inference workloads. And with that, I will hand over to Yuchen to, uh, walk you through more of it.

Agentic traces3:25

Yuchen Fama3:25

So we are currently, um, at this inflection point, moving from the era of classic LLM inference to the agentic era. So when we look at the real-world agentic workloads such as, uh, SweepBench and also WIKE traces from real-world cloud code sessions, they fundamentally break many assumptions we made with classic LLM serving.

Uh, as you heard actually many times in previous sessions, for example, multi-turns at new standard, we found from a few turns all the way to 3,000 turns. And also because agents frequently reuse the, uh, system prompt and the tool definitions, we usually see super high cache hit rate, um, oftentimes well exceeding 90%.

Uh, another thing is input-output ratio is, uh, is massive, oftentimes over 100:1 ratio and even higher in, you know, many cases. And on top of that, the context management is, is incredibly complex due to this high variance, because we can't just simply take the average and oftentimes we need to look at the distributions and the P90 numbers, especially when you do, uh, capacity planning.

And also, we observe really interesting patterns like sub-agent fan out, which is, which further complicates, uh, complicates scheduling. So to help communities study, um, these patterns, we collaborate with, uh, Google—thank you—and also IBM, our parent company, to add, uh, a trace replay tool in the inference perf.

You heard from earlier sessions, uh, from Ashok and Jason. Um, so yeah, feel free to check it out and the link is here.

Uh, next slide. Oh, so transition from the class, uh, the characteristics, um, we just saw. For agentic workloads, we're no longer chasing this, um, this, this raw throughput in a steady state. We often need to optimize, uh, for example, interactive latency, and they're very, um, highly volatile and client-driven contexts because user and, you know, client define the prompt structure.

Cache economics5:02

Yuchen Fama5:26

So this introduced several critical challenges. First of all, KV cache management becomes super volatile because the context is client-determined, as I said. So oftentimes we face this, like, you know, frequent evictions and rewrites. And secondly, we also need to tool, um, the engine like VOM with upper layer, uh, scheduling and routing.

It needs that coordination, such as prefix-aware routing, especially when latency becomes a primary, uh, scheduling matrix rather than like a secondary or afterthought. And thirdly, we also need to rethink our metrics. For example, we need to measure cache throughput separately.

Why? Because on theright, it's really clear that economic stakes is very high. So this is the, uh, Anthropic API pricing you also heard from earlier sessions. There's 10x cost difference between cache and non-cache tokens. So 10x difference on your, um, token balance sheet is, is pretty serious impact on your business.

So let's, let's look at how the KV cache is, um, both utilized and managed in LMD. So LMD router has this really flexible, um, endpoint picker plugins we call the EPP that can route the request to the optimal paths and that need the KV cache locality and also the load criteria.

Cache routing6:24

Yuchen Fama6:43

So the EPP continue probe each path, like VM pod metrics, to score each pod, um, like the running, for example, uh, running and waiting request. And then the KV cache utilization also prefix, uh, cache availability. And so we can schedule request to the optimal pod with the lowest load and also highest possibility to, um, to over cache hit.

So, um, going down from, uh, to the KV cache management layer, actually you also heard from, uh, earlier sessionright before this. So for agentic sessions when you have, uh, hot, warm, and cold cache, our current effort focuses on, for example, um, more offloading tiers like NVMe, SSD, and also, uh, file system XF, along with KV-centric store, um, like, uh, Mooncake.

And also implementing smarter and session-wide eviction policies such as priority and also session pinning to, uh, ensure this, uh, really important, you know, that the context persists exactly when and where it's needed.

So I'm going to play this, um, video really quick. Uh, it's a, it's a short demo.

Demo7:42

Ashish Kamra7:47

I don't want to stand here to look at it.

Yuchen Fama7:49

Okay.

So, okay. So this is an example of a KV cache-aware routing. As you see, when we send the very first request and it populate the KV cache, it takes roughly 3 seconds. And when we actually look at where it's, you know, the KV cache, uh, is going, there's no KV cache hit because it's the very first turn.

And then when we have the second turn, the request actually reuses the KV cache because, as you see, the system prompt is the same and this time takes about 1 second. And then when you actually look at the, uh, pod address, it's exactly the same because we define the KV cache.

Now going to the third turn, a new request with different system prompt. Now it takes about 3, uh, seconds. And, uh, as you see, you know,right now, and we don't find any KV cache hit because you can tell it's a different pod address.

And then if you just change the user prompt and keep the same system prompt, and the next turn you, you reuse the KV cache and then, then in this time it takes roughly about, uh, 1 second. Yeah, so it's a pretty intuitive demo.

And, um, I'll turn it to Ashish to talk about the next slide. But before that, what does the problem does it solve? So oftentimes the prefix-aware routing, KV cache routing helps you solve the TTFT problem. And of course, it will improve your latency, uh, your, your throughput.

But oftentimes for agentic workload, it's not just the TTFT. Your throughput is about your inter token latency. How do we solve that? So prefill, decode, disaggregation is a really, uh, powerful technique, but there are times that it works, there are times it doesn't work.

So I'll, uh, turn it to Ashish to give you a preview of, um, of the PD, uh, disaggregation.

Disaggregation9:29

Ashish Kamra9:29

So before we dive into PD, let's just, uh, look at what LLMD is. So LLMD is a high-performance Kubernetes native and actually now works on non-Kubernetes environments as well. Distributed LLM, LLM inference framework hosted under the CNCF umbrella.

LLMD provides a unified intelligent control plane designed specifically for agentic era of inference workloads. Well, Yuchen already talked about the router and the EPP at the top of the slide. Um, the other aspects are workload APIs such as leader-worker set and disaggregated set that orchestrates complex multi-multi-node model execution.

And then autoscalers that monitors capacity bounds and real-time traffic mixes to independently scale up and scale down, uh, your pods depending on the system load. So now look, now let's look at, uh, prefill, decode, disaggregation in detail. Um, uh, okay.

So why does PD exist in the first place? So one of the most powerful patterns implemented by LLMD is prefill, decode, disaggregation. And you must have heard from some of the previous talks as well. So what happens is in, in a non-PD situation in aggregated serving, one pod is responsible for optimizing both your time to first token and your inter token latencies.

Uh, but in PD, prefill and decode become independently scalable inference pods. But to understand why we actually need this, we have to look at, uh, the physics of LLM execution. So colocating, uh, both prefill and decode tasks on the same GPU creates something called as phase interference.

Prefill phase is the phase that creates the KV caches for your initial prompt. It wants high compute. It's highly bursty. Uh, utilizes GPUs at, uh, high flops and, and thrives on large batch parallelism to process the prompts and, uh, builds the initial KV cache.

The decode phase, on the other hand, is generating one token at a time and it's more mem-memory bandwidth hungry. It's highly latency sensitive and requires high KV cache residency. So in a, in a, in a traditional aggregated pod, if you, if there's a sudden influx of a long prefill prompt, it will completely stall the ongoing decode token generation process, causing massive, uh, problems and jitter in user streaming latency.

So, so how does PD actually work in practice in LLMD? So LLMD uses, um, uh, you know, like, okay, we'll start with step one. An incoming request hits the gateway router, which dynamically evaluates cluster states using something known as the endpoint picker, Yuchen talked about, and schedules the request to use PD disaggregation, selecting the optimal prefill and decode workers.

The router then coordinates the transaction directly with the designated prefill worker. The prefill worker processes the prompt, constructs the initial KV cache of the prompt, and outputs the standard KV transfer metadata. Um, and the target decode worker actually pulls the computed KV caches, um, uh, across the network fabric, fabric utilizing, uh, the KV transfer metadata that the, uh, uh, prefill pod had generated.

Um, okay. So with that, yes, that's kind of how, uh, PD is implemented in practice in LLMD. And next, I would like to show you some, uh, experimental results on where PD actually shines. So in this graph, you can see that, um, uh, in, in, in the standard aggregated deployment, which is the top red line, uh, the P99 ITL, uh, hovers roughly around 900 milliseconds.

PD results12:50

Ashish Kamra13:17

And you can, you can see some fluctuations, um, up and down. And, but the, the bottom blue line is the P99, uh, inter token latency on a PD deployment. And you can see that it's drastically almost 9 times better at 100 milliseconds.

And it's also much smoother, uh, than the aggregated serving.

And, uh, this is some of our own internal results at Red Hat. So for a gpt-oss 120B model, uh, 16 H100s, uh, the aggregated config is, uh, four replicas, uh, Tensor Parallelism 4. And the disaggregated is two prefill, two decode, all with Tensor Parallelism 4.

It's a highly multi-turn workload with a 10,000 token prefix and 128 tokens for every turn, uh, every turn. So, so this is a great chart. Like you can see at the bottom most line is a standard aggregated config that, uh, is doing the default Kubernetes scheduling and, uh, and it's aggregated.

So that's kind of our baseline. And then the middle blue line is still aggregated, but with the LLMD, uh, KV cache-aware routing. And you can almost see the gains just, just based on the routing. And the red line is actually the PD, uh, the prefill decode config with two prefill and two decode workers.

And you can actually see that, like, it's very similar to the aggregated config at the lower concurrency regimes and, uh, even and, and very similar at the higher concurrency regimes. But it's, it's actually the middle part of the concurrency regime that PD actually shines.

And, and these are some of the, the classic pareto curves that we see when you actually do PD and, uh, aggregated side by side. So these results are again from the gpt-oss 120B model. 64 H100s aggregated is eight replicas, TP8, and disag is, uh, three prefill, five decode, again TP8.

And, uh, prefill heavy workload with like 5,000 average input sequence length and 500 output sequence length. And you can actually see the blue line is the, the PD curve and the red line is the aggregated curve. And the PD curve kind of dominates, um, uh, the aggregate curve across the entire interactivity spectrum.

PD tradeoffs15:37

Ashish Kamra15:37

Okay. But I don't want to leave you guys that PD is the answer to everything and it's a magic bullet. But, um, it's, uh, it's essentially a separation phase separation trade-off and not a magic bullet. So we created this, uh, matrix to help you decide when PD might be, uh, good for you.

So if you're managing long context, uh, with high ISL-OSL ratios, and you if you have a large model that you're serving that can, that you can apply rich model parallelism techniques, um, you're facing that middle concurrency regime, uh, that I, I showed you in the previous graphs.

And, and the very important part is that if you want, uh, strict ITL streaming requirements, like you want the, you want the token generation to be, uh, much more smooth, um, then you want to consider PD. But we also saw that it requires transfer of KV caches from your prefill workers to your decode workers.

So you must pro-process an advanced, uh, high-speed network fabric like, uh, RDMA or, uh, RoCE to support that KV cache transfer. And if you do not have such requirements, short, moderate context, any model size, low, uh, low concurrency regimes, or, uh, if you have strict TTFT requirements because you can actually tune them on an aggregate serving, um, and you the biggest point is like if you don't have the network fabric to support those KV cache transfers, so you might actually just want to stick with aggregated.

So here is my key takeaway from all of this. So architecting this complex platform requires balancing a lot of, uh, knobs and a highly multi-dimensional design space, all of which is supported in LLMD as you saw. The scheduler must support or constantly evaluate SLO targets, uh, queue depths, KV cache locality metrics, PD ratios, and network topologies to be able to route the request to the optimal pod.

While in, while the PD design space, you, you need dynamic PD rate matching to adapt to PD ratios because, you know, you can start with a static PD ratio, but it needs to evolve with the autoscaler as the traffic changes.

Um, and you need, uh, yeah, autoscaling to scale PD pools independently, um, and constantly tweaking model parallelism techniques like Tensor Parallelism, Data Parallelism, uh, to meet your SLOs. So, um, I think with these, uh, I will hand it over to Yuchen to anchor some of the concepts that we showed the real-world case study of serving the GLM 5.2 model, uh, which is, uh, still ongoing as we speak.

Yuchen Fama18:12

Yeah, still ongoing. You probably have seen tons of, uh, impressive numbers of GLM 5.2 on B200. When we talk to our customers and they usually don't have, you know, the luxury of B200, they have a lot of H200s.

Case study18:12

Yuchen Fama18:25

So we have to figure out how to, like, put all the knobs together and make GLM 5.2 work really well for cluster of, of H200s. So, uh, we let's anchor all the concept together. Um, we went through, for example, the, uh, KV cache routing, PD disaggregation.

We kind of call them a wildlift path in LLMD. And also we combined with different parallelism strategies too, so we can, uh, independently, uh, scale prefill pods because for agentic workload is super, uh, long, you know, like heavy prefill.

So, uh, in this case, we designed the prefill pool using, uh, up to three workers, optimized for, uh, high throughput, uh, with deep EP. And then for decode pool, we use, uh, one dedicated worker and, um, that's optimized for, for low latency.

So we use Nixel for efficient KV transfer between the pools. And, uh, also with the each worker, we have the leader worker set group, uh, with TP1, DP8, and also, uh, EP8, uh, Expert Parallelism 8. So the architecture is just highly modular because you can, uh, actually scale the throughput by simply adding, uh, prefill workers without reconfiguring and, um, the decode pool.

So, uh, this highlights how LLMD effectively, effectively manage the complexity of combining like TP and DP and EP at scale. And also we found some interesting fun fact actually a couple of days ago. Um, BF6, uh, BF16 KV cache actually is faster than using like FPA, uh, KV cache for longer prefill.

Um, this is also like we continue to explore and found like more interesting patterns. But more importantly, uh, we want to kind of just show the result really quick. So, um, for this, uh, data set, agentic workload data set, the ISL-OSL ratio is pretty high, 45 to 1 ratio.

Prefill is, uh, is really the constraint you can tell. Um, with 2P, even 1D, we have, um, 4x faster TTFT and also 60, uh, percent more requests. And this is continuous like work in progress. So the next step is we need to also tune the upper layer scheduler, lower TTFT, and also adding more, more prefill replicas.

Future20:26

Yuchen Fama20:26

So, um, I know we are running out of time really quick. Uh, we, um, the fundamental shift for agentic workload, we're continuing to, uh, have this LLMD, uh, agentic North, uh, North Star, uh, with session graph orchestration, program war scheduling, uh, state reuse, lifecycle, and also the, uh, agentic benchmark, um, we're working on.

So you can find them, uh, in LLMD, ups, uh, upstream LLMD. And also, you know, feel free to join the SIG group and, uh, and contribute. And, um, this is the very last slide. So distributed inference is not a challenge, uh, every single comp a single company can solve alone.

We're proud to be, uh, building this, uh, future in the open alongside our incredible ecosystem collaborators, uh, CoreWeave, Google, IBM, NVIDIA, Growing List launch partners, and industry adopters. So if you're passionate about the future of open source inference, we invite you to join us.

We do have a booth downstairs. Feel free to stop by, ask any questions. And, uh, thank you so much for your time.