AI-Dynamo published Dynamo v1.6.0-deepseek-v4.1-flash-dev.1 on September 12, an experimental snapshot build that serves DeepSeek-V4.1-Flash on the Dynamo SGLang backend with up to 1,048,576 tokens of context1. The release includes Kubernetes recipes for 8x GB200 aggregated deployment with KV-aware routing and DSpark speculative decoding, and a 1P1D disaggregated profile using Mooncake KV transfer over TCP or GKE RDMA. The build ships native Rust prompt rendering, reasoning and tool-call parsing, and runs on CUDA 13, but is not QA-gated and carries no published performance benchmarks.
Dynamo v1.6.0 Adds DeepSeek-V4.1-Flash Serving on GB200 With 1M-Token Context
AI-Dynamo's experimental v1.6.0 snapshot build adds DeepSeek-V4.1-Flash serving on the SGLang backend with 1M-token context and Kubernetes recipes for…