🔬Study Finds Repair Hallucinations in 72.7% of LLM Patches
TL;DR
Manual review of 812 LLM-generated patches found repair hallucinations in 72.7% of cases, including patches that passed every available test. Only 21.0% to 55.9% of generated patches cleared the developer-written suite across three models.
Manual review of 812 LLM-generated patches found repair hallucinations in 72.7% of cases, including patches that passed every available test. Only 21.0% to 55.9% of generated patches cleared the developer-written suite across three models.
Key Points
arXiv 2609.04909, submitted Sep 4 by Xuemeng Cai and four co-authors
Three representative LLMs evaluated on 832 Defects4J bugs
21.0%-55.9% of patches pass the developer-written test suite
Incorrect causal localization drives 45.9% of hallucinations; wrong repair strategy 18.5%
Models also misidentify triggering test cases and mispredict line coverage on branches
Why It Matters
A green test run is not evidence your agent understood the bug, so gate auto-fixes on causal localization rather than on CI going green.
Quick Facts
Frequently Asked Questions
Why does this matter?
A green test run is not evidence your agent understood the bug, so gate auto-fixes on causal localization rather than on CI going green.
What happened?
Manual review of 812 LLM-generated patches found repair hallucinations in 72.7% of cases, including patches that passed every available test. Only 21.0% to 55.9% of generated patches cleared the developer-written suite across three models.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,463 builders reading daily.