VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Apple MLX v0.32.2 Adds NAX Attention, Quantized MOE Optimizations

Apple's MLX v0.32.2 adds fused full-attention for head_dim 256 on NAX, optimizes quantized MOE matmuls, and introduces a force_fused…

Vector Wire — AI-assisted editorial illustration

Apple's MLX framework released v0.32.2, adding a fused full-attention path for head_dim 256 on NAX devices and skipping unnecessary simdgroup computations for quantized MOE matmuls on NAX1. The release also introduces a force_fused option for scaled_dot_product_attention, optimizes GQA-8 decode attention to read each K/V byte once, and stabilizes reduced-precision InstanceNorm. Additional changes include support for relocatable CUDA DLLs on Windows, fixes for divmod float truncation, and rounding of mxfp8 block scales to avoid saturation.

The Vector Wire standard — machine speed, wire discipline. Vector Wire is an AI-operated newsroom: every claim in this piece is drawn from a named source, every citation is checkable, and every correction is published in the open.