Skip to content
theregister·

💡AMD Partners With Cerebras for Ultra-Low Latency Inference

Ultra-fast inference just got more accessible

TL;DR

AMD and Cerebras are partnering to deliver ultra-low-latency inference, aiming for a 5x boost in tokens per watt. The combined offering will be available on Cerebras Cloud later this year.

AMD has teamed up with Cerebras Systems to develop an ultra-low-latency disaggregated compute platform aimed at agentic workloads. This collaboration promises to significantly enhance interactivity and efficiency, delivering a 5x increase in tokens per watt of electricity consumed. For developers working on large-scale AI models like Kimi K2.5, this partnership means needing just a few dozen accelerators instead of thousands, drastically reducing costs and complexity. The combined offering will be available on Cerebras Cloud later this year.

AMD Partners With Cerebras for Ultra-Low Latency Inference — theregister

Key Points

1

Cerebras' wafer scale engines (WSE) use on-chip SRAM that's orders of magnitude faster than HBM4

2

The partnership aims to achieve higher interactivity without compromising throughput or cost

3

AMD and Cerebras' solution is expected to boost the number of tokens per second generated per watt by up to 5x

4

Cerebras Cloud will offer this combined offering later in 2023, targeting agentic workloads

5

The collaboration was announced during an Advancing AI keynote on Thursday

Why It Matters

If you're working with large-scale AI models like Kimi K2.5, AMD and Cerebras' solution means needing just a few dozen accelerators instead of thousands, drastically reducing costs and complexity. This is particularly relevant for teams focused on ultra-low-latency inference in agentic workloads.

AMDCerebrasAI InferenceDisaggregated Compute

Frequently Asked Questions

Why does this matter?

If you're working with large-scale AI models like Kimi K2.5, AMD and Cerebras' solution means needing just a few dozen accelerators instead of thousands, drastically reducing costs and complexity. This is particularly relevant for teams focused on ultra-low-latency inference in agentic workloads.

What happened?

AMD and Cerebras are partnering to deliver ultra-low-latency inference, aiming for a 5x boost in tokens per watt. The combined offering will be available on Cerebras Cloud later this year.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 2,228 builders reading daily.