Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

AMD FP8 Training Optimizations Merged Into PyTorch TorchTitan and TorchAO

AMD and Meta engineers upstreamed FP8 training optimizations into TorchTitan and TorchAO, delivering 13.4% throughput gains on Llama3-8B and up to 6.2×…

AMD and Meta/PyTorch engineers have upstreamed FP8 training optimizations into pytorch/AO and pytorch/TorchTitan, enabling AMD Instinct GPU support with competitive FP8 performance out of the box1. On Llama3-8B, FP8 training delivers a 13.4% throughput gain over BF16. For DeepSeek-V3 671B MoE shapes, fused Triton quantization kernels recovered 89% of FP8 quantization overhead, with individual kernel optimizations delivering up to a 6.2× speedup. At the PyTorch Conference 2025, the team demonstrated linear scaling beyond 1,000 GPUs on AMD Instinct clusters.