Skip to content
theregister·

💡Cerebras Unveils WSE-3T With Double Compute and Memory

Cerebras' new chip doubles compute, but...

TL;DR

Cerebras unveiled the WSE-3T, doubling its compute and memory. While impressive, it relies heavily on sparsity for peak performance, limiting benefits for LLM inference. AWS and AMD partnerships offload some workload.

Cerebras just dropped the WSE-3T, promising double the compute power of its predecessor. But here's the thing: while this chip has a whopping 250 petaFLOPS of AI compute and 44 GB of SRAM, it relies on sparsity for peak performance, which isn't as beneficial for LLM inference. The real story is Cerebras' partnerships with AWS and AMD to offload some workload onto their respective accelerators. If you're working with large datasets or heavy computations, this could change your workflow.

Cerebras Unveils WSE-3T With Double Compute and Memory — theregister

Key Points

1

WSE-3T delivers 250 petaFLOPS of AI compute, up from 128 in WSE-3.

2

Memory bandwidth jumps to 43.2 PB/s with 44 GB SRAM onboard.

3

New system uses 3x the accelerators per rack, consuming around 140 kW total.

4

Cerebras partnered with AWS and AMD for offloading compute-intensive tasks.

5

The WSE-3T cuts latency from five to two microseconds via pipeline parallelism.

Why It Matters

If you're running large-scale AI workloads or heavy computations, the WSE-3T's double the compute power could be a game-changer. However, its reliance on sparsity for peak performance means it might not be as beneficial for LLM inference. The partnerships with AWS and AMD offer potential cost savings by offloading some workload.

CerebrasWSE-3TAI ComputeSparsity

Frequently Asked Questions

Why does this matter?

If you're running large-scale AI workloads or heavy computations, the WSE-3T's double the compute power could be a game-changer. However, its reliance on sparsity for peak performance means it might not be as beneficial for LLM inference. The partnerships with AWS and AMD offer potential cost savings by offloading some workload.

What happened?

Cerebras unveiled the WSE-3T, doubling its compute and memory. While impressive, it relies heavily on sparsity for peak performance, limiting benefits for LLM inference. AWS and AMD partnerships offload some workload.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,180 builders reading daily.

Also get