🔬AI Agents Reproduced 2,226 ICML Papers, Falsified 496
TL;DR
Hugging Face ran a 19-day hackathon where 1,221 people pointed coding agents at ICML 2026 papers claim by claim. They published 6,816 logbooks covering 2,226 papers, about a third of the conference.
Hugging Face ran a 19-day hackathon where 1,221 people pointed coding agents at ICML 2026 papers claim by claim. They published 6,816 logbooks covering 2,226 papers, about a third of the conference. 51% of examined papers had a claim verified; 23% had one falsified or contested.
Key Points
6,816 Trackio logbooks, 35,908 claims judged, 2,962 HF Jobs launched on $20 credit budgets
266 papers fully reproduced; 49 had every claim falsified
242 papers where independent teams reached opposite verdicts on the same claims
A spotlight paper's paging guarantee was corrected from Hk+O(1) to Hk+Θ(log k), confirmed at ~9 sigma
A QKV-variants paper's headline figure moved from 3.1% to ~9.4% quality cost once EOS padding was excluded
Why It Matters
The failure modes are the takeaway: agents verified the broken paper because their checks stopped before log-k growth appeared, so horizon length is now a first-class variable in any agentic eval you write.
Quick Facts
Frequently Asked Questions
Why does this matter?
The failure modes are the takeaway: agents verified the broken paper because their checks stopped before log-k growth appeared, so horizon length is now a first-class variable in any agentic eval you write.
What happened?
Hugging Face ran a 19-day hackathon where 1,221 people pointed coding agents at ICML 2026 papers claim by claim. They published 6,816 logbooks covering 2,226 papers, about a third of the conference.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,313 builders reading daily.