🔬Google: LLMs Store 95-98% of Facts But Fail to Recall Them
TL;DR
Google Research finds frontier models encode 95-98% of tested facts in their parameters yet fail to recall 26-34% of them on demand. Longer thinking recovers 40-65% of those buried facts.
Google Research finds frontier models encode 95-98% of tested facts in their parameters yet fail to recall 26-34% of them on demand. Longer thinking recovers 40-65% of those buried facts. The bottleneck has moved from what models know to what they can actually retrieve.

Key Points
Evaluated on Gemini 2.5 Pro, Gemini 3 Pro, Gemini 3 Flash and GPT-5
95-98% of facts encoded; 26-34% not directly recalled
Extended thinking recovers 40-65% of encoded-but-unrecalled facts
Even with thinking, 11-12% of facts still fail
Rare facts are encoded nearly as often as popular ones; the gap sits in recall
Why It Matters
If your RAG stack exists to patch missing knowledge, you may be paying to fix the wrong failure. Better retrieval prompting and a larger reasoning budget can beat adding another index.
Quick Facts
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,484 builders reading daily.