VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

Helion Kernels Ship via Hugging Face Hub, Outpace SDPA and Flash-Linear-Attention

Meta's Helion DSL gains Hugging Face Kernels support, shipping attention kernels 1.20x faster than SDPA and linear-attention kernels 1.41x faster than…

Meta's Helion, a high-level DSL for writing portable ML kernels, is now supported within Hugging Face's Kernels project, enabling developers to package, autotune, and distribute kernels through the Hugging Face Hub1. A shipped attention kernel outperforms PyTorch SDPA on 19 of 19 pre-tuned shapes with a 1.20x geomean speedup on Nvidia H100s. Seven linear-attention kernels, pre-tuned for Nvidia B200s, beat flash-linear-attention across all variants with a 1.41x geomean device-time speedup on pre-tuned shapes and 1.35x on held-out shapes.