Skip to content
daily-hour-news·

Arm's CSS for Mobile 2 Targets On-Device Agentic AI

TL;DR

Arm announced Compute Subsystem for Mobile 2 in Shanghai, built on the argument that agentic workloads need several fast CPUs at once. It doubles SME2 matrix units to two and adds the first Mali GPU with neural acceleration, claiming 70% faster small language models.

Arm announced Compute Subsystem for Mobile 2 in Shanghai, built on the argument that agentic workloads need several fast CPUs at once. It doubles SME2 matrix units to two and adds the first Mali GPU with neural acceleration, claiming 70% faster small language models.

Arm's CSS for Mobile 2 Targets On-Device Agentic AI — daily-hour-news

Key Points

1

Flagship cluster: two C2-Ultra cores, six C2-Pro efficiency cores and two SME2 units, up from one SME2 unit in last year's Lumex CSS

2

Arm claims C2-Ultra delivers up to 15% higher single-thread performance than C1-Ultra, and doubled SME2 gives up to 70% speedup on small language models

3

Mali G2-Ultra NX is Arm's first AI-native GPU, with Neural Super Sampling upscaling 540p to 1080p and Neural Frame Rate Upscaling generating intermediate frames

4

Combined, the GPU renders one eighth of displayed pixels and reconstructs the rest, targeting 30fps ray-traced scenes inside a 1 W budget

5

Neural graphics SDK and Vulkan sample code are on GitHub; silicon from licensees is expected in devices next year at the earliest

Why It Matters

Arm is betting that agent orchestration moves to the handset, so the mobile constraint shifts from peak single-core speed to keeping several concurrent inference workloads responsive.

Quick Facts

ArmCSS for Mobile 2edge AIAI agentsMali GPUSME2on-device inference

Frequently Asked Questions

Why does this matter?

Arm is betting that agent orchestration moves to the handset, so the mobile constraint shifts from peak single-core speed to keeping several concurrent inference workloads responsive.

What happened?

Arm announced Compute Subsystem for Mobile 2 in Shanghai, built on the argument that agentic workloads need several fast CPUs at once. It doubles SME2 matrix units to two and adds the first Mali GPU with neural acceleration, claiming 70% faster small language models.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,467 builders reading daily.

Also get