💻Google Splits TPU 8 into Training and Inference Modes
Your cloud bill is about to change
TL;DR
Google Cloud is splitting its eighth-generation TPU into two modes: TPU 8t for model training and TPU 8i for inference. The new chips are up to 3x faster and offer 80% better performance per dollar compared to previous generations.
Google just split its custom-built AI chip, the TPU 8, into two modes: TPU 8t for model training and TPU 8i for inference. This means you'll get up to 3x faster training times and 80% better performance per dollar compared to previous generations. The new chips can scale to support over 1 million TPUs in a single cluster, making them ideal for large-scale AI workloads.

Key Points
The new TPU 8t mode is up to 3x faster for AI model training compared to previous generations
The new TPUs offer 80% better performance per dollar compared to previous generations
Google Cloud's custom-built chips can scale to support over 1 million TPUs in a single cluster
Nvidia is not being replaced, but rather supplemented with Google's own custom-built chips
Google Cloud will offer Nvidia's latest chip, Vera Rubin, in its infrastructure later this year
Why It Matters
If you're running large-scale AI workloads on AWS or Azure, this changes the economics. The new TPUs are designed to provide more compute power at lower energy and cost, making them ideal for high-performance computing. However, it's worth noting that Nvidia is still a dominant player in the market, and its $5 trillion market cap isn't going anywhere anytime soon.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.