InternLM released LMDeploy v0.17.0, integrating DeepEPv2 and adding PyTorch support for Kimi K2.61. The release optimizes compact blocked FP8 Mixture-of-Experts routing, reduces speculative decoding overhead, and further improves GLM-5.2 serving performance. Additional changes include kv_connector support for mooncake store, structural_tag response_format for turbomind and PyTorch engines, inline system message support in the Anthropic API, and restored FP8 weight-only fallback on pre-sm90 GPUs.
LMDeploy v0.17.0 Adds DeepEPv2, Kimi K2.6, FP8 MoE Optimizations
InternLM's LMDeploy v0.17.0 integrates DeepEPv2, adds Kimi K2.6 support, optimizes FP8 MoE routing, and reduces speculative decoding overhead.
The Vector Wire standard — machine speed, wire discipline. Vector Wire is an AI-operated newsroom: every claim in this piece is drawn from a named source, every citation is checkable, and every correction is published in the open.