🔬ThinkingBox: 67% of Failed Agent Runs Still Say Done
TL;DR
Microsoft's ThinkingBox ran 121,680 trials across 12 models on 507 stateful workflows. In 67.24% of failed runs the agent ended cleanly with no error.
Microsoft's ThinkingBox ran 121,680 trials across 12 models on 507 stateful workflows. In 67.24% of failed runs the agent ended cleanly with no error. Claude Opus 5.5 led at 67.16% pass@1 but held just 47.53% across 20 repeats.

Key Points
121,680 trials, 12 models, 507 workflows graded on final database state
67.24% of failed checks still terminated cleanly with valid tool calls
Kimi-K3 solves 93.89% of tasks at least once but only 13.41% on all 20 tries
Cost per dependable task: GPT-5.4 $6.80, GPT-6 Astra $7.45, Claude Opus 5.5 $7.80
79.9% of failures come from tool handling and error recovery, not reasoning
Why It Matters
Pass@1 flatters agents. If you ship one to production, measure repeat reliability and end-state correctness, and invest in retry logic before a bigger model.
Quick Facts
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,556 builders reading daily.