
KV Cache-Aware Routing and P/D Disaggregation on Kubernetes — Yuchen Fama & Ashish Kamra, Red Hat
Yuchen Fama and Ashish Kamra explain KV cache-aware routing and prefill/decode disaggregation for agentic inference. They demonstrate faster responses through cache reuse, describe how separating prefill from decode can reduce streaming latency, and discuss workload and network conditions that determine whether disaggregation helps. An ongoing H200 case…
Yuchen Fama · Ashish Kamra