Skip to content
daily-hour-news·

🔬ThinkingBox: 67% of Failed Agent Runs Still Say Done

TL;DR

Microsoft's ThinkingBox ran 121,680 trials across 12 models on 507 stateful workflows. In 67.24% of failed runs the agent ended cleanly with no error.

Microsoft's ThinkingBox ran 121,680 trials across 12 models on 507 stateful workflows. In 67.24% of failed runs the agent ended cleanly with no error. Claude Opus 5.5 led at 67.16% pass@1 but held just 47.53% across 20 repeats.

ThinkingBox: 67% of Failed Agent Runs Still Say Done — daily-hour-news

Key Points

1

121,680 trials, 12 models, 507 workflows graded on final database state

2

67.24% of failed checks still terminated cleanly with valid tool calls

3

Kimi-K3 solves 93.89% of tasks at least once but only 13.41% on all 20 tries

4

Cost per dependable task: GPT-5.4 $6.80, GPT-6 Astra $7.45, Claude Opus 5.5 $7.80

5

79.9% of failures come from tool handling and error recovery, not reasoning

Why It Matters

Pass@1 flatters agents. If you ship one to production, measure repeat reliability and end-state correctness, and invest in retry logic before a bigger model.

Quick Facts

agent evaluationThinkingBoxMicrosoftbenchmarkreliabilityClaude Opus 5.5tool usepass@k

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,556 builders reading daily.