🔬Prime Agent Lifts ARC-AGI-3 Best@1 From 30% to 95.5%
TL;DR
Prime Intellect's open-source harness pairs a persistent IPython REPL with memory, skills and subagent specs that survive across trajectories. On ARC-AGI-3 RHAE it takes Best@1 from 30% to 95.5%.
Prime Intellect's open-source harness pairs a persistent IPython REPL with memory, skills and subagent specs that survive across trajectories. On ARC-AGI-3 RHAE it takes Best@1 from 30% to 95.5%. The argument: most 'model failures' on long tasks are harness failures.

Key Points
Posted to arXiv on August 24, 2026 by Prime Intellect
A persistent IPython REPL implements the Recursive Language Model abstraction for programmatic context handling and test-time compute
Continual Harness carries histories, memories, skills, prompts and subagent specifications between runs
Recursive subagents talk directly to each other; an Agents View lets humans inspect daemon-backed sessions
Matches or beats native harnesses on long-context coding, GPU kernel generation, emulator construction and autonomous nanoGPT speedruns
On Factorio, dedicated subagents enabled parallel work and continuous tech progression
Why It Matters
If a harness swing is worth 65 points on a reasoning benchmark, published model rankings are measuring scaffolding as much as weights. Budget your engineering accordingly.
Quick Facts
Frequently Asked Questions
Why does this matter?
If a harness swing is worth 65 points on a reasoning benchmark, published model rankings are measuring scaffolding as much as weights. Budget your engineering accordingly.
What happened?
Prime Intellect's open-source harness pairs a persistent IPython REPL with memory, skills and subagent specs that survive across trajectories. On ARC-AGI-3 RHAE it takes Best@1 from 30% to 95.5%.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,317 builders reading daily.