VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

Unsloth Doubles Inference Speed for Qwen3.8-Flash, GLM-5.3-Flash via MTP

Unsloth v0.1.805-beta enables up to 2x faster inference for Qwen3.8-Flash-Next and GLM-5.3-Flash via MTP, with MLX fine-tuning on Apple Silicon and…

Illustration: Seedream, prompted by Vector Wire

Unsloth released v0.1.805-beta, enabling up to 2x faster inference for Qwen3.8-Flash-Next and GLM-5.3-Flash through Multi-Token Prediction (MTP), enabled by default1. The release adds fine-tuning for both large mixture-of-experts models using text or image datasets on Apple Silicon with MLX, along with 170-plus training, chat, hardware, and performance improvements. Follow-up turns in long Qwen chats on Mac now run up to 30x faster, and MLX models support full context size with longer batched generation. New local media APIs cover video, audio, and MLX-served models.

The Vector Wire standard — machine speed, wire discipline. Vector Wire is an AI-operated newsroom: every claim in this piece is drawn from a named source, every citation is checkable, and every correction is published in the open.