Skip to content
daily-hour-news·

🔬DeepWeb-Bench Finds Retrieval Isn't the Research Bottleneck

TL;DR

A new arXiv benchmark, DeepWeb-Bench, tested nine frontier models on open-web deep research. Retrieval caused only 12-14% of errors, while derivation and calibration drove over 70%, with strong and weak models failing in different ways.

A new arXiv benchmark, DeepWeb-Bench, tested nine frontier models on open-web deep research. Retrieval caused only 12-14% of errors, while derivation and calibration drove over 70%, with strong and weak models failing in different ways.

DeepWeb-Bench Finds Retrieval Isn't the Research Bottleneck — daily-hour-news

Key Points

1

Posted to arXiv on May 20, 2026; nine frontier models evaluated

2

Retrieval failures account for just 12-14% of errors

3

Derivation and calibration failures cause over 70% of errors

4

Strong models stumble on incomplete derivation; weak ones hallucinate precision

5

Cross-model agreement was only rho = 0.61, showing real domain specialization

Why It Matters

If search isn't the weak link, teams building research agents should spend their effort on reasoning and calibration, not on bigger retrieval pipelines.

Quick Facts

DeepWeb-Benchresearch agentsbenchmarkarXivdeep researchevaluation

Frequently Asked Questions

Why does this matter?

If search isn't the weak link, teams building research agents should spend their effort on reasoning and calibration, not on bigger retrieval pipelines.

What happened?

A new arXiv benchmark, DeepWeb-Bench, tested nine frontier models on open-web deep research. Retrieval caused only 12-14% of errors, while derivation and calibration drove over 70%, with strong and weak models failing in different ways.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,094 builders reading daily.

Also get