Skip to content
daily-hour-news·

🦾RobotWorld: Frontier Models Fail 63 of 84 Robot Tasks

TL;DR

A new Hugging Face benchmark runs five frontier multimodal agents on 84 simulated robot tasks. None solved 63 of them, and the best model, GPT-6 Astra, cleared 16.

A new Hugging Face benchmark runs five frontier multimodal agents on 84 simulated robot tasks. None solved 63 of them, and the best model, GPT-6 Astra, cleared 16.

RobotWorld: Frontier Models Fail 63 of 84 Robot Tasks — daily-hour-news

Key Points

1

Tasks span manipulation, mobile manipulation, locomotion, driving and aerial control

2

GPT-6 Astra solved 16 of 84, Claude Opus 5.5 solved 13, Kimi K3 solved 2

3

Opus 5.5 cleared 3 of 4 aerial-control tasks; Astra led on manipulation at 23.7%

4

Authors say the bottleneck is composing capabilities, not perception

5

Astra's evaluation cost about $9,913 across 941M tokens

Why It Matters

General-purpose models are still far from running robots by prompt alone. Humanoid demos look smooth; this is the cold shower for anyone betting on a VLM as the robot brain.

Quick Facts

roboticsbenchmarkvision-language modelsGPT-6 AstraClaude Opus 5.5embodied AIHugging Face

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,558 builders reading daily.