NVIDIA's AI Dynamo project released v1.4.0-qwen-3.8-2.4t-dev.1, an experimental snapshot build adding serving support for Qwen/Qwen3.8-2.4T-A95B-FP8 on both vLLM and SGLang backends1. The build, committed August 27, includes Kubernetes recipes for vLLM aggregate chat and agentic profiles and SGLang aggregate and disaggregated chat profiles on GB300 and GB200 hardware using TP16 MNNVL configurations. The vLLM runtime is built on vLLM 0.27.1 with CUDA 13.0. NVIDIA labeled the release not production-ready and not QA-gated, intended for evaluation and early feedback only.
NVIDIA Dynamo Adds Qwen 2.4T Serving on GB300, GB200
NVIDIA's Dynamo v1.4.0 experimental build adds serving for Qwen3.8-2.4T-A95B-FP8 on vLLM and SGLang backends with Kubernetes recipes for GB300 and GB200…
The Vector Wire standard — machine speed, wire discipline. Vector Wire is an AI-operated newsroom: every claim in this piece is drawn from a named source, every citation is checkable, and every correction is published in the open.