💡Cerebras Unveils WSE-3T Chip With 250 PetaFLOPS AI Compute
New Cerebras chip doubles compute, cuts latency by half
TL;DR
Cerebras' new WSE-3T chip offers 250 petaFLOPS and 44 GB SRAM, doubling compute and cutting latency. AWS and AMD partnerships optimize performance.
Cerebras just unveiled the WSE-3T, a next-gen AI accelerator that doubles its predecessor's compute power to 250 petaFLOPS. This chip is crucial for teams working with massive models like LLMs, offering unprecedented memory bandwidth and reduced latency. The new system supports up to three 'backpack' form factors per rack, pushing total system power to around 140 kW. With AWS and AMD partnerships offloading compute-intensive tasks, the WSE-3T could be a game-changer for large-scale AI projects.

Key Points
WSE-3T delivers 250 petaFLOPS AI compute and 44 GB SRAM, up from WSE-3's 128 petaFLOPS and 24 GB SRAM.
Each CS-4 system can be equipped with up to three 'backpack' form factors for additional performance.
Cerebras partnered with AWS Trainium XPUs and AMD Instinct GPUs to optimize compute-intensive tasks in the inference pipeline.
Power consumption increased to around 120 kW - 140 kW per rack, but power delivery is twice as efficient enabling higher operating frequencies.
Latency reduced from five microseconds down to two thanks to improved chip-to-chip bandwidth and direct interconnects.
Why It Matters
If you're running large-scale AI models like LLMs in production, the WSE-3T's 250 petaFLOPS compute power and reduced latency can significantly improve performance. However, the system's high power consumption means it might not be suitable for all environments.
Frequently Asked Questions
Why does this matter?
If you're running large-scale AI models like LLMs in production, the WSE-3T's 250 petaFLOPS compute power and reduced latency can significantly improve performance. However, the system's high power consumption means it might not be suitable for all environments.
What happened?
Cerebras' new WSE-3T chip offers 250 petaFLOPS and 44 GB SRAM, doubling compute and cutting latency. AWS and AMD partnerships optimize performance.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,180 builders reading daily.