Skip to content
MIT Technology Review·

💡AI Solves NYT Puzzles Near Perfectly in 2025

AI's puzzle-solving skills are closing the gap with humans

TL;DR

AI models can now solve New York Times Connections puzzles near perfectly, up from just 18% in late 2024. But they still struggle with visual and spatial reasoning tasks.

AI models have made a leap in solving New York Times Connections puzzles, going from a 18% success rate in late 2024 to near-perfect solutions today. This showcases the rapid improvement in AI's problem-solving abilities but also highlights its limitations in handling visual and spatial reasoning tasks. For instance, models still struggle with classic riddles and visual puzzles, where humans have a clear edge. As the complexity of puzzles increases, such as in the Tower of Hanoi or river-crossing scenarios, AI models begin to falter when the number of disks or people hits six and higher. This indicates that while AI excels in pattern recognition and memorization, it still lacks the nuanced understanding and adaptability of human spatial reasoning.

AI Solves NYT Puzzles Near Perfectly in 2025 — MIT Technology Review

Key Points

1

In late 2024, AI models could solve only 18% of NYT Connections puzzles, but by early 2025, that number shot up to near-perfect accuracy.

2

Models still trip up on subtle changes in classic riddles and visual puzzles, where humans excel in spatial reasoning.

3

LLMs struggle with logic grid puzzles when the number of individuals or clues exceeds a certain threshold, typically around six.

4

ARC-AGI questions, which require inferring abstract rules, still stump top-tier models when presented as images rather than strings of numbers.

5

Models show remarkable memory capacity, having been trained on vast datasets, but often rely on memorized patterns rather than generalizable rules.

Why It Matters

If you're working on AI-driven problem-solving tools, the rapid improvement in puzzle-solving skills is a red flag. Models may appear to be getting smarter, but they're still prone to overfitting on training data and failing on novel tasks. This impacts the reliability of AI in real-world applications where adaptability is key.

aipuzzlesproblem-solvinglimitationsprogress

Frequently Asked Questions

Why does this matter?

If you're working on AI-driven problem-solving tools, the rapid improvement in puzzle-solving skills is a red flag. Models may appear to be getting smarter, but they're still prone to overfitting on training data and failing on novel tasks. This impacts the reliability of AI in real-world applications where adaptability is key.

What happened?

AI models can now solve New York Times Connections puzzles near perfectly, up from just 18% in late 2024. But they still struggle with visual and spatial reasoning tasks.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,316 builders reading daily.

Also get