Meta disclosed a Triton-based attention kernel called Jagged Flash Attention (JFA), built with Triton Low-level Extensions (TLX), that outperforms FlashAttention-4 (May 2026 version) on NVIDIA B200 GPUs for the variable-length sequences used in Meta's Generative Ads Model (GEM)1. On jagged shapes in bfloat16, the kernel is approximately 13% faster on the forward pass and 50% faster on the backward pass than FA4, while requiring roughly 3.2K lines of code compared to FA4's approximately 10K-line CuteDSL implementation. Attention is the single slowest kernel in GEM, and TLX closes the gap between high-level Triton and hand-written CUDA by exposing explicit SMEM/TMEM allocation, warp specialization, and async TMA as first-class primitives. Code is available on GitHub under Meta's ads_model_kernel_library repository.
Meta's TLX Attention Kernel Outperforms FlashAttention-4 on Blackwell GPUs
Meta's TLX-based Jagged Flash Attention kernel beats FlashAttention-4 by 13% forward and 50% backward on NVIDIA B200 GPUs in 3x less code.