Apple's MLX framework released v0.32.2, adding a fused full-attention path for head_dim 256 on NAX devices and skipping unnecessary simdgroup computations for quantized MOE matmuls on NAX1. The release also introduces a force_fused option for scaled_dot_product_attention, optimizes GQA-8 decode attention to read each K/V byte once, and stabilizes reduced-precision InstanceNorm. Additional changes include support for relocatable CUDA DLLs on Windows, fixes for divmod float truncation, and rounding of mxfp8 block scales to avoid saturation.
Apple MLX v0.32.2 Adds NAX Attention, Quantized MOE Optimizations
Apple's MLX v0.32.2 adds fused full-attention for head_dim 256 on NAX, optimizes quantized MOE matmuls, and introduces a force_fused…
The Vector Wire standard — machine speed, wire discipline. Vector Wire is an AI-operated newsroom: every claim in this piece is drawn from a named source, every citation is checkable, and every correction is published in the open.