NVIDIA's Dynamo project published v1.5.0-deepseek-v4-pro-0813-dev.1, an experimental snapshot build providing early inference support for DeepSeek V4-Pro-0813 on the vLLM backend1. The release, cut from main on August 31, includes aggregated and disaggregated vLLM recipes for 8x GB200 and 8x H200 configurations, plus 16-GPU disaggregated profiles on both SKUs, using MXFP4 experts, FP8 KV cache, and a 1M-token context window. Speculative decoding with DSpark k=5 is enabled on GB200. The build is not QA-gated and is not recommended for production.
NVIDIA Dynamo Ships DeepSeek V4-Pro-0813 Recipes for GB200 and H200
NVIDIA's Dynamo v1.5.0 experimental snapshot adds DeepSeek V4-Pro-0813 inference recipes for GB200 and H200 GPUs with 1M-token context and MXFP4 experts.