Skip to content
MIT Technology Review·

💡AI Agents Fail NeurIPS Research Test

AI can't do open-ended research yet

TL;DR

A study testing AI agents on open-ended research tasks found they struggle with creativity and judgment, despite excelling at engineering. Anthropic's Claude Opus 4.8 was given six days and $3k API credits but failed to meet publication standards.

AI researchers set out to test whether advanced models could conduct original scientific research. They tasked AI agents like Anthropic's Claude with producing papers for NeurIPS, a top-tier conference. The agents managed the engineering tasks but fell short on creativity and judgment, failing to incorporate feedback or refine their hypotheses effectively. This study highlights that while AI can automate many aspects of research, it still lacks the human touch necessary for groundbreaking discoveries. The results suggest recursive self-improvement may be further off than some predict.

AI Agents Fail NeurIPS Research Test — MIT Technology Review

Key Points

1

Researchers gave Claude Opus 4.8 six days and $3k API credits to write a paper for NeurIPS 2026, which it failed to produce

2

The study used 'shadow evaluation' to assess agents' ability to complete tasks without human oversight, finding they struggle with open-ended thinking

3

Claude developed novel hypotheses but rejected them due to limited data and struggled to incorporate feedback from reviewing tools

4

Anthropic's blog post on the experiment details its progress toward models that speed up their own development through reinforcement learning

5

The study suggests AI is good at narrow tasks but lacks creativity for open-ended research, tempering claims of imminent recursive self-improvement

Why It Matters

If you're working with advanced AI agents on complex research projects, this study shows they can handle the technical aspects but lack human judgment and creativity. This means relying solely on AI for groundbreaking research isn't feasible yet.

AIResearchOpen-ended thinkingReinforcement learningNeurIPS

Frequently Asked Questions

Why does this matter?

If you're working with advanced AI agents on complex research projects, this study shows they can handle the technical aspects but lack human judgment and creativity. This means relying solely on AI for groundbreaking research isn't feasible yet.

What happened?

A study testing AI agents on open-ended research tasks found they struggle with creativity and judgment, despite excelling at engineering. Anthropic's Claude Opus 4.8 was given six days and $3k API credits but failed to meet publication standards.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,155 builders reading daily.

Also get