Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Home

PyTorch's FBTriton Kernels Outperform CUDA on Embedding Workloads

PyTorch's FBTriton Triton kernels achieve a median 1.28x forward speedup over legacy CUDA across 307 shard configurations for Table Batched Embedding…

Meta's PyTorch team said its FBTriton Triton-based kernels for Table Batched Embedding operations outperform legacy CUDA kernels across recommendation-system workloads, according to a PyTorch blog post1. Across 307 shard configurations, the median forward speedup is 1.28x. On a B200 GPU configuration, combined forward-backward latency dropped from 79.537 ms to 66.183 ms. TBE kernels handle embedding lookups across thousands of sharded GPUs, combining lookup and pooling in a single GPU launch to reduce overhead.