🔬New Benchmark Shows Agent Memory Fails in Real Chats
TL;DR
LOCOMO-CONV tests five agent memory systems on how people actually talk, not on QA-style probes. Retrieval gaps show up on implicit and composed queries that QA benchmarks miss entirely.
LOCOMO-CONV tests five agent memory systems on how people actually talk, not on QA-style probes. Retrieval gaps show up on implicit and composed queries that QA benchmarks miss entirely. Strong retrieval also fails to translate into better responses.
Key Points
Four query styles derived from LoCoMo: dialog, implicit, counterfactual, and composed
Measures retrieval recall and end-to-end response quality across five memory systems
Multi-facet query rewriting narrows the gap for raw-turn memory but not abstractive memory
Documents silent grounding, where memory improves answers without surfacing the gold fact
Authors Wen-Yu Chang and Yun-Nung Chen release supportive_memory annotations with the benchmark
Why It Matters
If your agent memory looks healthy on QA evals, this paper argues you are measuring the wrong thing and shipping retrieval holes.
Quick Facts
Frequently Asked Questions
Why does this matter?
If your agent memory looks healthy on QA evals, this paper argues you are measuring the wrong thing and shipping retrieval holes.
What happened?
LOCOMO-CONV tests five agent memory systems on how people actually talk, not on QA-style probes. Retrieval gaps show up on implicit and composed queries that QA benchmarks miss entirely.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,461 builders reading daily.