VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

AMD FP8 Training Optimizations Merged Into PyTorch TorchTitan and TorchAO

AMD and Meta engineers upstreamed FP8 training optimizations into TorchTitan and TorchAO, delivering 13.4% throughput gains on Llama3-8B and up to 6.2×…

Vector Wire — AI-assisted editorial illustration

AMD and Meta/PyTorch engineers have upstreamed FP8 training optimizations into pytorch/AO and pytorch/TorchTitan, enabling AMD Instinct GPU support with competitive FP8 performance out of the box1. On Llama3-8B, FP8 training delivers a 13.4% throughput gain over BF16. For DeepSeek-V3 671B MoE shapes, fused Triton quantization kernels recovered 89% of FP8 quantization overhead, with individual kernel optimizations delivering up to a 6.2× speedup. At the PyTorch Conference 2025, the team demonstrated linear scaling beyond 1,000 GPUs on AMD Instinct clusters.

CORRECTIONS: none for this article · this piece updates automatically as the story develops · corrections policy & trail →