🤖Meta's MTIA 400 Chip: 12 PetaFLOPS for AI Training
Meta's new chip is a beast for AI training, but it's not all it's cracked up to be
TL;DR
Meta's MTIA 400 chip delivers 12 petaFLOPS for AI training, but it's 3x slower than rivals. It's aimed at LLM training and will be followed by a faster inference version next year.
Meta's MTIA 400 chip is designed for AI training and ad serving, delivering 12 petaFLOPS of MXFP4 compute. It's 20% faster than Nvidia's top Blackwell accelerators at higher precisions but falls short compared to AMD and Nvidia's latest chips. The chip's architecture includes two compute dies, two I/O dies, and an SoC die, built using Broadcom's XPU tech. Each compute die features a 6x8 grid of processing elements on a 3nm process, and the chip is fed by eight 36 GB HBM3e stacks, delivering 288 GB of memory and 9.2 TB/s of bandwidth. A single rack is equipped with 18 compute blades and eight switch blades, totaling 72 accelerators. Meta is developing an inference-optimized version, the MTIA 450, set for production next year, with the MTIA 500 expected in 2027.

Key Points
Meta's MTIA 400 chip delivers 12 petaFLOPS of MXFP4 compute at 1.7 GHz.
The chip features two compute dies, two I/O dies, and an SoC die, built using Broadcom's XPU tech.
Each compute die has a 6x8 grid of processing elements on a 3nm process, delivering 288 GB of memory and 9.2 TB/s of bandwidth.
A single rack is equipped with 18 compute blades and eight switch blades, totaling 72 accelerators.
Meta is developing an inference-optimized version, the MTIA 450, set for production next year, with the MTIA 500 expected in 2027.
Why It Matters
If you're training large language models (LLMs), Meta's MTIA 400 chip offers 12 petaFLOPS of compute, but it's 3x slower than the latest Nvidia and AMD chips. The MTIA 450, set for production next year, will double the memory bandwidth, making it a more competitive option for inference workloads. However, the initial performance gap means teams should carefully evaluate if the MTIA 400 is the right choice for their current needs.
Frequently Asked Questions
Why does this matter?
If you're training large language models (LLMs), Meta's MTIA 400 chip offers 12 petaFLOPS of compute, but it's 3x slower than the latest Nvidia and AMD chips. The MTIA 450, set for production next year, will double the memory bandwidth, making it a more competitive option for inference workloads. However, the initial performance gap means teams should carefully evaluate if the MTIA 400 is the right choice for their current needs.
What happened?
Meta's MTIA 400 chip delivers 12 petaFLOPS for AI training, but it's 3x slower than rivals. It's aimed at LLM training and will be followed by a faster inference version next year.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,337 builders reading daily.