🛡️OpenAI Pauses Astra at 'Critical' Cyber Threshold
TL;DR
OpenAI paused internal work on its Astra model after evaluations suggested it may reach the 'Critical' cyber tier of its Preparedness Framework. That rating means a model could find and run zero-day exploits against hardened systems without human help.
OpenAI paused internal work on its Astra model after evaluations suggested it may reach the 'Critical' cyber tier of its Preparedness Framework. That rating means a model could find and run zero-day exploits against hardened systems without human help.

Key Points
OpenAI cannot yet rule out Critical cyber capability for Astra after recent internal evals
Critical tier = autonomous discovery and execution of zero-day exploits on hardened targets
Company added universal monitoring for risky actions across all agentic uses of Astra
OpenAI says it will test the model with government agencies and outside safety orgs
Why It Matters
A frontier lab voluntarily halting a model at its own danger line is the first real test of whether preparedness frameworks change behavior when the stakes are high.
Quick Facts
Frequently Asked Questions
Why does this matter?
A frontier lab voluntarily halting a model at its own danger line is the first real test of whether preparedness frameworks change behavior when the stakes are high.
What happened?
OpenAI paused internal work on its Astra model after evaluations suggested it may reach the 'Critical' cyber tier of its Preparedness Framework. That rating means a model could find and run zero-day exploits against hardened systems without human help.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 2,763 builders reading daily.