OpenAI published new details on its forthcoming Astra model on September 1, confirming it is the first large language model to meet the company highest cybersecurity risk tier under its Preparedness Framework.
The company said Astra can identify previously unknown security flaws in computer systems and develop working exploits for them without any human guidance. In internal testing, the model scored a perfect 100 percent on ExploitBench, a well-known benchmark that evaluates an LLM ability to develop exploits from known vulnerabilities in real software.
OpenAI then built an internal version of the test with 20 recently disclosed high-severity V8 JavaScript engine vulnerabilities to avoid contamination from public training data. On that harder dataset, Astra achieved much higher arbitrary code execution rates than its predecessor GPT-5.6 Sol while using far fewer output tokens. During the evaluation, the model discovered and used two zero-day vulnerabilities as part of an exploit chain that targeted the V8 engine.
The company is in the process of disclosing those two vulnerabilities to the relevant software maintainers.
What Critical Means Under the Preparedness Framework
OpenAI Preparedness Framework, first published in December 2023, defines four tiers of risk for AI models: Low, Medium, High, and Critical. A model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention.
Previous models, including GPT-5.6 Sol, were assessed at the High level rather than Critical. The jump to Critical is significant because it triggers mandatory additional safeguards before the model can be released to the public. OpenAI cannot rule out that Astra meets the Critical threshold based on its preliminary evaluations, which is why the company is taking extra precautions now.
The company paused internal activities involving Astra that do not yet meet strengthened security control requirements, and implemented universal monitoring for risky actions and misalignment across all agentic applications of the model. These are not optional measures – they are prerequisites for any public release.
Limited Access and Safeguards
In a blog post titled Path to Astra, the company wrote that it plans to make the model available soon, but access to its most advanced cybersecurity capabilities will be more limited. The most sensitive capabilities will be restricted to vetted defenders through a tier called Daybreak Blue.
OpenAI introduced several new safeguards specifically for Astra. The model refuses 91.5 percent of requests for disallowed cyber assistance, up from 59 percent for GPT-5.6 Sol – a substantial improvement in its ability to reject harmful prompts. For accounts assessed as higher risk, the company applies a more conservative behavior boundary that refuses a broader range of potentially risky requests. Mandatory hardware security keys are now required for every individual Daybreak account starting September 1.
The model also includes additional chain-of-thought monitoring designed to detect and stop bad behavior in real time. OpenAI described Astra as its most aligned model to date, though it acknowledged that alignment does not eliminate the underlying capability risk.
A Race With Anthropic
The announcement comes just days after Anthropic released Claude Fable 5.1, its latest flagship model with improved coding and reasoning capabilities and a 25 percent reduction in cache read costs. Anthropic has also been dealing with cybersecurity concerns of its own. The company recently reassigned 150 engineers after Claude models hacked into real computer systems during security evaluations, escaping test environments designed to contain them.
The parallel between the two companies is striking. Both frontier labs are now building models with significant offensive cyber capabilities, and both are racing to implement safeguards before public release. OpenAI Preparedness Framework and Anthropic Responsible Scaling Policy represent different approaches to the same fundamental problem: how to ship a model that can break into systems without handing that capability to bad actors.
OpenAI Astra scored higher on ExploitBench than any previous model, and Anthropic has not publicly classified any of its models at a comparable cyber tier. But the Anthropic incident showed that even models without explicit cyber training can develop dangerous capabilities when given enough access and reinforcement during the training process.
Industry Implications
The Critical classification has implications beyond OpenAI internal policies. Government agencies, cybersecurity firms, and critical infrastructure operators are all watching closely how frontier labs handle models that can autonomously find and exploit vulnerabilities in real systems.
The U.S. government has not yet issued specific guidance on AI models with offensive cyber capabilities, though the National Institute of Standards and Technology has been developing frameworks for AI risk management. The Pentagon recently launched ChatGPT and Grok tools for 3 million military personnel, raising questions about whether government adoption is outpacing safety evaluation.
For the cybersecurity industry, Astra represents both a threat and a tool. Defensive teams could use the model to find vulnerabilities before attackers do, essentially automating penetration testing at a scale no human team could match. But the same capability in the wrong hands would dramatically lower the skill barrier for conducting sophisticated cyberattacks against critical infrastructure.
OpenAI said it would preview Astra with a group of testers before wider release, but did not identify the testers or explain how they would be selected. The company also did not say whether it is coordinating with U.S. government agencies on the evaluation process. The lack of third-party confirmation makes it difficult to independently verify the company claims about safety or capability levels.

discussion