TL;DR: Agentic workloads generate massive shared prefixes that are often recomputed across turns. By integrating Mooncake’s distributed KV cache store into vLLM, we achieve 3.8x higher throughput,…
Serving Agentic Workloads at Scale with vLLM x Mooncake
calendar_today
May 6, 2026
person
Yifan Qiao, Trong Dao Le, Ao Shen, Zhewen Li, Bowen Wang
domain
vllm