🦾RobotWorld: Frontier Models Fail 63 of 84 Robot Tasks
TL;DR
A new Hugging Face benchmark runs five frontier multimodal agents on 84 simulated robot tasks. None solved 63 of them, and the best model, GPT-6 Astra, cleared 16.
A new Hugging Face benchmark runs five frontier multimodal agents on 84 simulated robot tasks. None solved 63 of them, and the best model, GPT-6 Astra, cleared 16.
Key Points
Tasks span manipulation, mobile manipulation, locomotion, driving and aerial control
GPT-6 Astra solved 16 of 84, Claude Opus 5.5 solved 13, Kimi K3 solved 2
Opus 5.5 cleared 3 of 4 aerial-control tasks; Astra led on manipulation at 23.7%
Authors say the bottleneck is composing capabilities, not perception
Astra's evaluation cost about $9,913 across 941M tokens
Why It Matters
General-purpose models are still far from running robots by prompt alone. Humanoid demos look smooth; this is the cold shower for anyone betting on a VLM as the robot brain.
Quick Facts
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,558 builders reading daily.