💡AMD Partners With Cerebras for Ultra-Low Latency Inference
Ultra-fast inference just got more accessible
TL;DR
AMD and Cerebras are partnering to deliver ultra-low-latency inference, aiming for a 5x boost in tokens per watt. The combined offering will be available on Cerebras Cloud later this year.
AMD has teamed up with Cerebras Systems to develop an ultra-low-latency disaggregated compute platform aimed at agentic workloads. This collaboration promises to significantly enhance interactivity and efficiency, delivering a 5x increase in tokens per watt of electricity consumed. For developers working on large-scale AI models like Kimi K2.5, this partnership means needing just a few dozen accelerators instead of thousands, drastically reducing costs and complexity. The combined offering will be available on Cerebras Cloud later this year.

Key Points
Cerebras' wafer scale engines (WSE) use on-chip SRAM that's orders of magnitude faster than HBM4
The partnership aims to achieve higher interactivity without compromising throughput or cost
AMD and Cerebras' solution is expected to boost the number of tokens per second generated per watt by up to 5x
Cerebras Cloud will offer this combined offering later in 2023, targeting agentic workloads
The collaboration was announced during an Advancing AI keynote on Thursday
Why It Matters
If you're working with large-scale AI models like Kimi K2.5, AMD and Cerebras' solution means needing just a few dozen accelerators instead of thousands, drastically reducing costs and complexity. This is particularly relevant for teams focused on ultra-low-latency inference in agentic workloads.
Frequently Asked Questions
Why does this matter?
If you're working with large-scale AI models like Kimi K2.5, AMD and Cerebras' solution means needing just a few dozen accelerators instead of thousands, drastically reducing costs and complexity. This is particularly relevant for teams focused on ultra-low-latency inference in agentic workloads.
What happened?
AMD and Cerebras are partnering to deliver ultra-low-latency inference, aiming for a 5x boost in tokens per watt. The combined offering will be available on Cerebras Cloud later this year.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 2,228 builders reading daily.