Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Home

NVIDIA Dynamo Adds Session-Aware Routing for Agentic Inference Across vLLM, SGLang

NVIDIA Dynamo uses a session-level identifier to enable session-aware routing and KV cache reuse across vLLM and SGLang, achieving over 94.5%…

NVIDIA Dynamo introduced a unified session-level identifier that converts request-level inference serving into a program-aware system, enabling session-aware routing, shared KV cache indexing, and programmatic cache movement across vLLM and SGLang1. On a Uni-Agent SWE-Bench sweep using two TP4 MiniMax-M2 replicas on a single 8xH100 node, Dynamo held prefix-cache hit rates above 94.5% and achieved 11.0–14.6% higher model-token throughput than VERL's default Global LB at medium concurrency. Claude Code, Codex, and OpenCode work with Dynamo without additional configuration.