Skip to content
TechCrunch·

💻Google Splits TPU 8 into Training and Inference Modes

Your cloud bill is about to change

TL;DR

Google Cloud is splitting its eighth-generation TPU into two modes: TPU 8t for model training and TPU 8i for inference. The new chips are up to 3x faster and offer 80% better performance per dollar compared to previous generations.

Google just split its custom-built AI chip, the TPU 8, into two modes: TPU 8t for model training and TPU 8i for inference. This means you'll get up to 3x faster training times and 80% better performance per dollar compared to previous generations. The new chips can scale to support over 1 million TPUs in a single cluster, making them ideal for large-scale AI workloads.

Google Splits TPU 8 into Training and Inference Modes — TechCrunch

Key Points

1

The new TPU 8t mode is up to 3x faster for AI model training compared to previous generations

2

The new TPUs offer 80% better performance per dollar compared to previous generations

3

Google Cloud's custom-built chips can scale to support over 1 million TPUs in a single cluster

4

Nvidia is not being replaced, but rather supplemented with Google's own custom-built chips

5

Google Cloud will offer Nvidia's latest chip, Vera Rubin, in its infrastructure later this year

Why It Matters

If you're running large-scale AI workloads on AWS or Azure, this changes the economics. The new TPUs are designed to provide more compute power at lower energy and cost, making them ideal for high-performance computing. However, it's worth noting that Nvidia is still a dominant player in the market, and its $5 trillion market cap isn't going anywhere anytime soon.

Google CloudNvidiaTPUAI chips

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.