🔬DeepWeb-Bench Finds Retrieval Isn't the Research Bottleneck
TL;DR
A new arXiv benchmark, DeepWeb-Bench, tested nine frontier models on open-web deep research. Retrieval caused only 12-14% of errors, while derivation and calibration drove over 70%, with strong and weak models failing in different ways.
A new arXiv benchmark, DeepWeb-Bench, tested nine frontier models on open-web deep research. Retrieval caused only 12-14% of errors, while derivation and calibration drove over 70%, with strong and weak models failing in different ways.

Key Points
Posted to arXiv on May 20, 2026; nine frontier models evaluated
Retrieval failures account for just 12-14% of errors
Derivation and calibration failures cause over 70% of errors
Strong models stumble on incomplete derivation; weak ones hallucinate precision
Cross-model agreement was only rho = 0.61, showing real domain specialization
Why It Matters
If search isn't the weak link, teams building research agents should spend their effort on reasoning and calibration, not on bigger retrieval pipelines.
Quick Facts
Frequently Asked Questions
Why does this matter?
If search isn't the weak link, teams building research agents should spend their effort on reasoning and calibration, not on bigger retrieval pipelines.
What happened?
A new arXiv benchmark, DeepWeb-Bench, tested nine frontier models on open-web deep research. Retrieval caused only 12-14% of errors, while derivation and calibration drove over 70%, with strong and weak models failing in different ways.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,303 builders reading daily.