Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Helion Backend Outperforms CUTLASS, DeepGEMM in vLLM Inference on Hopper GPUs

Meta's Helion kernel DSL integrated into vLLM outperforms CUTLASS and DeepGEMM on NVIDIA Hopper GPUs, achieving up to 1.178× speedups across quantized…

Meta's Helion kernel DSL, integrated into vLLM's linear backend, outperforms default CUTLASS and DeepGEMM backends across evaluated models on NVIDIA Hopper GPUs, delivering more than 10% throughput improvement for some workloads, according to a PyTorch blog post by contributors from Red Hat and Meta1. A single Helion GEMM implementation covers Standard GEMM, Split-K, and Swap-AB variants, with per-shape autotuning selecting the best configuration. Benchmarks on an NVIDIA H100 80GB GPU showed geometric mean speedups of 1.177× for Block_FP8 over DeepGEMM, 1.178× for W8A8_INT8 over CUTLASS, and 1.110× for FP8_Dynamic over CUTLASS.