AMD and Meta/PyTorch engineers have upstreamed FP8 training optimizations into pytorch/AO and pytorch/TorchTitan, enabling AMD Instinct GPU support with competitive FP8 performance out of the box1. On Llama3-8B, FP8 training delivers a 13.4% throughput gain over BF16. For DeepSeek-V3 671B MoE shapes, fused Triton quantization kernels recovered 89% of FP8 quantization overhead, with individual kernel optimizations delivering up to a 6.2× speedup. At the PyTorch Conference 2025, the team demonstrated linear scaling beyond 1,000 GPUs on AMD Instinct clusters.
AMD FP8 Training Optimizations Merged Into PyTorch TorchTitan and TorchAO
AMD and Meta engineers upstreamed FP8 training optimizations into TorchTitan and TorchAO, delivering 13.4% throughput gains on Llama3-8B and up to 6.2×…
CORRECTIONS: none for this article · this piece updates automatically as the story develops · corrections policy & trail →