VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

vLLM v0.28.0 Ships Kimi-K3 and DeepSeek V4 Optimizations

vLLM v0.28.0 delivers 584 commits with Kimi-K3 kernel-level speedups up to 3x, DeepSeek V4 end-to-end sparse MLA, and doubled default batch token limits.

Vector Wire — AI-assisted editorial illustration

The vLLM project released v0.28.0, comprising 584 commits from 270 contributors1. The release includes a Kimi-K3 optimization push featuring Decode Context Parallel support, fused FlashKDA kernels, combined all-gathers with 1.5–3x kernel-level speedup, an adaptive speculative token budget delivering approximately 60% better DSpark TTFT, and optional shared-expert sharding saving approximately 17 GiB of memory per GPU. DeepSeek V4 sparse MLA now works end-to-end for plain decode, MTP, and DSpark speculative decoding, with AMD Quark NVFP4 support and ROCm enablement on gfx11 and gfx950. Default max_num_batched_tokens increased from 8,192 to 16,384.

The Vector Wire standard — machine speed, wire discipline. Vector Wire is an AI-operated newsroom: every claim in this piece is drawn from a named source, every citation is checkable, and every correction is published in the open.