Skip to content
ARC Prize·

🤖GPT-6 Astra Surpasses Human Baseline on ARC-AGI-3

AI outperforms humans in complex reasoning tasks

TL;DR

GPT-6 Astra scores 99.9% on ARC-AGI-3 with Provider Adapter, surpassing human baseline. It uses fewer actions and costs less, showing AI's efficiency in complex reasoning tasks.

GPT-6 Astra has scored 99.9% on the ARC-AGI-3 benchmark, surpassing the human baseline in action efficiency. This matters for developers working on AI-driven decision-making systems, as it demonstrates the model's ability to solve complex tasks more efficiently than humans. GPT-6 Astra uses fewer actions and costs less, with scores of 62.7% for $26K and 99.9% for $19K. The model's compact algebraic notation and custom tools are key to its efficiency. Developers should watch how this impacts AI-driven decision-making systems.

GPT-6 Astra Surpasses Human Baseline on ARC-AGI-3 — ARC Prize

Key Points

1

GPT-6 Astra scores 62.7% on ARC-AGI-3 with Standard harness for $26K.

2

GPT-6 Astra scores 99.9% on ARC-AGI-3 with Provider Adapter harness for $19K.

3

GPT-6 Astra uses fewer actions than median human on 96% of levels.

4

Human participants were paid $115 per 90-minute session, plus $5 per game.

5

GPT-6 Astra's compact algebraic notation and custom tools boost efficiency.

Why It Matters

If you're building AI-driven decision-making systems, GPT-6 Astra's performance on ARC-AGI-3 shows it can solve complex tasks more efficiently than humans. The model's compact notation and custom tools reduce costs and improve efficiency, making it a strong choice for developers looking to optimize AI-driven workflows. However, the $19K cost for 99.9% performance may be prohibitive for smaller projects.

GPT-6 AstraARC-AGI-3AI efficiencycomplex reasoningdecision-making systems

Frequently Asked Questions

Why does this matter?

If you're building AI-driven decision-making systems, GPT-6 Astra's performance on ARC-AGI-3 shows it can solve complex tasks more efficiently than humans. The model's compact notation and custom tools reduce costs and improve efficiency, making it a strong choice for developers looking to optimize AI-driven workflows. However, the $19K cost for 99.9% performance may be prohibitive for smaller projects.

What happened?

GPT-6 Astra scores 99.9% on ARC-AGI-3 with Provider Adapter, surpassing human baseline. It uses fewer actions and costs less, showing AI's efficiency in complex reasoning tasks.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,465 builders reading daily.

Also get