Meta disclosed MTIA 300, the first in its family of custom training and inference accelerators designed for ranking and recommendation models1. The chip integrates two network chiplets containing twelve custom 800 Gbps RDMA NICs, delivering 1.2 TB/s of total I/O bandwidth without crossing a PCIe bus. Sixteen dedicated message engines, each with a RISC-V core and near-memory compute block, handle communication independently, enabling line-rate AllReduce and ReduceScatter collectives without touching the compute grid. Meta reported that running large GEMMs concurrently with collective operations introduces less than 0.5% compute degradation, compared to over 20% on traditional GPUs. On a 150-billion-parameter production recommendation model across 40 accelerators, MTIA 300 communication time was 3.9 times faster than an equivalent GPU cluster.
Meta Unveils MTIA 300, Its First In-House AI Training Chip with Built-in NICs
Meta disclosed MTIA 300, its first custom AI training chip with integrated NIC chiplets delivering 1.2 TB/s I/O bandwidth and 3.9x faster communication…