OpenAI’s Astra AI Model Becomes First to Hit ‘Critical’ Cybersecurity Rating

OpenAI's Astra AI Model Becomes First to Hit 'Critical' Cybersecurity Rating

OpenAI says its Astra model is the first AI system to reach the Critical cybersecurity threshold under its Preparedness Framework, triggering tighter controls before any release.

Imagine an AI system that can independently find hidden security flaws in a hospital’s computer network, a power grid’s control systems, or a government database — and then work out how to exploit them, all without a human guiding each step. That’s roughly what OpenAI is describing when it says its upcoming Astra model has crossed a threshold it calls “Critical” in cybersecurity capability. And it’s the first time any of the company’s models has done so.

OpenAI posted about the development as part of a broader communication effort explaining how Astra was assessed under the company’s Preparedness Framework, stating that the company is “focused on making increasingly capable AI safe and broadly accessible.” But the designation has raised serious questions across the technology and security communities about what it means to release a model with these kinds of capabilities at all.

What Is the Preparedness Framework?

OpenAI introduced its Preparedness Framework back in 2023 and updated it in 2025. Think of it as a risk-management rulebook for the company’s most powerful AI models. It tracks capabilities across several categories — cybersecurity, biological and chemical risks, and AI self-improvement — and assigns one of two main warning levels: High, meaning a model can amplify existing pathways to severe harm, and Critical, meaning a model can open up entirely new pathways to severe harm that didn’t exist before.

That distinction matters. High is serious. Critical is something else entirely.

For cybersecurity specifically, a model reaches the Critical threshold if it can either identify and develop working exploits of previously unknown security vulnerabilities — so-called zero-day flaws — across many well-protected real-world systems without human intervention, or devise and carry out end-to-end cyberattack strategies against hardened targets when given nothing more than a high-level goal.

Previous OpenAI models, including o1 and gpt-oss-120b, were publicly documented as not reaching even the High threshold in cybersecurity. Astra is the first to reach Critical.

How Did OpenAI Get Here?

The story moved in stages. On 7 August 2026, OpenAI publicly acknowledged that internal evaluations of Astra showed it had made significant advances in agentic coding and cybersecurity — meaning it could act more independently across complex tasks — and that the company “could not rule out” Critical cyber capability. At that point, OpenAI tightened controls on Astra-related work and paused certain internal activities that didn’t meet its strengthened security standards.

By late August and into early September 2026, subsequent blog posts and external reporting confirmed that OpenAI had formally decided to treat Astra as meeting the Critical cybersecurity threshold. It’s a subtle but meaningful shift: from “we can’t rule it out” to “it does.”

OpenAI has said Astra will be released — but only once the company judges that its safeguards sufficiently reduce the risk of severe harm. No precise release date or commercial pricing has been confirmed publicly.

The Safeguards OpenAI Is Putting in Place

Because of the Critical designation, OpenAI says it has introduced stronger monitoring of training and inference workloads involving tools, along with tighter limits on certain internal uses of the model. The company has framed Astra’s development as a test case for how the industry should handle AI models that approach or cross these capability thresholds — including slowing or constraining development until adequate protections are in place.

But not everyone is convinced that’s enough.

Critics and digital-rights advocates have argued that releasing any model with Critical cyber capability, even with safeguards, introduces systemic risk to global digital infrastructure — above all if access controls fail or the model is somehow replicated or leaked. Some have called for independent external audits rather than relying on OpenAI’s own internal assessments, along with clearer public benchmarks and firmer commitments about what would actually delay or limit a release.

There are also concerns about how OpenAI has communicated the shift from “cannot rule out Critical” to “is Critical”, with some commentators suggesting the language blurs the line between uncertainty and confirmed risk.

Cybersecurity professionals, for their part, tend to see both sides. A model capable of independently finding zero-day vulnerabilities could be a powerful tool for defensive security work — helping organisations discover and patch flaws before attackers do. But the same capability in the wrong hands, or with inadequate access controls, could supercharge the threat landscape considerably.

Sam Altman, OpenAI’s chief executive, has previously said that the company believes safety and capability can advance together, though he has not made specific public comments tied directly to Astra’s Critical designation at the time of writing.

What This Means for Kent Residents

For people here in Kent, the practical impact of Astra’s Critical designation is indirect but real. Public bodies that rely on digital systems — Kent County Council, NHS Kent and Medway Integrated Care Board, Kent Police, and local schools — could face more sophisticated cyberattacks if AI models with these capabilities were ever misused or accessed by malicious actors. On the other side, defensive cybersecurity teams working with local authorities and NHS networks could potentially use similar capabilities to find and fix vulnerabilities before attackers do. The National Cyber Security Centre is likely to factor developments like Astra into its national guidance, which filters down into the technical standards and procurement policies that Kent organisations are expected to follow — so while residents won’t feel this overnight, it will shape the digital security environment around them.

Source: @OpenAI

⚡

OpenAI's Astra AI Model Becomes First to Hit 'Critical' Cybersecurity Rating Quiz

5 questions