Skip to content
daily-hour-news·

🔬Google: LLMs Store 95-98% of Facts But Fail to Recall Them

TL;DR

Google Research finds frontier models encode 95-98% of tested facts in their parameters yet fail to recall 26-34% of them on demand. Longer thinking recovers 40-65% of those buried facts.

Google Research finds frontier models encode 95-98% of tested facts in their parameters yet fail to recall 26-34% of them on demand. Longer thinking recovers 40-65% of those buried facts. The bottleneck has moved from what models know to what they can actually retrieve.

Google: LLMs Store 95-98% of Facts But Fail to Recall Them — daily-hour-news

Key Points

1

Evaluated on Gemini 2.5 Pro, Gemini 3 Pro, Gemini 3 Flash and GPT-5

2

95-98% of facts encoded; 26-34% not directly recalled

3

Extended thinking recovers 40-65% of encoded-but-unrecalled facts

4

Even with thinking, 11-12% of facts still fail

5

Rare facts are encoded nearly as often as popular ones; the gap sits in recall

Why It Matters

If your RAG stack exists to patch missing knowledge, you may be paying to fix the wrong failure. Better retrieval prompting and a larger reasoning budget can beat adding another index.

Quick Facts

Google ResearchfactualityhallucinationRAGGeminiGPT-5evaluation

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,484 builders reading daily.

Also get