OpenAI's Astra Becomes First AI Flagged as Critically Dangerous Before Release
OpenAI's upcoming Astra model has crossed a new internal safety threshold for cybersecurity, triggering mandatory development controls and a new era for AI safety governance.

- OpenAI has classified its upcoming model Astra as potentially "Critical" for cybersecurity under its Preparedness Framework — a first for any OpenAI model.
- "Critical" means a model may autonomously find and exploit zero-day vulnerabilities in hardened real-world systems with no human direction — a higher bar than any prior model including GPT-5.6-Sol.
- OpenAI is pausing internal Astra activities that don't meet new security requirements and moving development into isolated, sandboxed environments with universal Chain-of-Thought monitoring.
- Astra was not involved in the Hugging Face sandbox escape incident; that involved separate evaluation models running with reduced safety restrictions.
- OpenAI has voluntarily notified the White House and will partner with government agencies and AI safety organizations to test Astra's capabilities before any release.
- The long-term goal remains deploying Astra through the Daybreak program to verified defenders — but only once containment infrastructure is proven robust.
OpenAI has declared its upcoming model, Astra, the first AI system it has ever classified as potentially "Critical" for cybersecurity under its Preparedness Framework , the company's internal risk governance policy. This is not a story about a model being shelved. It is a story about a safety framework doing exactly what it was designed to do, and about what it means that we have now reached the capability level it was written to anticipate.
What "Critical" actually means
The Preparedness Framework, first published in late 2023 and substantially revised in April 2025, defines two risk thresholds for frontier models: High and Critical. The distinction matters enormously.
- High: A model at "High" capability could amplify existing pathways to severe harm and must have safeguards that sufficiently reduce that risk before it is deployed.
- Critical: A model at "Critical" capability could open unprecedented new pathways to harm and must have safeguards in place not only before deployment but during development.
In practical terms, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high-level desired goal. That is a qualitatively different kind of threat than anything OpenAI has previously flagged.
Every prior model, including GPT-5.6-Sol, sat a level below, at "High". Astra is the first to potentially cross into Critical territory.
What the evals actually found
OpenAI's latest internal evaluations of Astra over the past few days indicate significant advancements in agentic coding and cybersecurity. These results, in addition to expert assessments, led OpenAI to conclude that it cannot rule out critical cyber capabilities under its Preparedness Framework.
The language here is deliberate and precise. OpenAI is not saying Astra