IT Brief US - Technology news for CIOs & IT decision-makers
United States
OpenAI says Astra may hit critical cyber threshold

OpenAI says Astra may hit critical cyber threshold

Mon, 10th Aug 2026 (Today)
Mark Tarre
MARK TARRE News Chief

OpenAI has said one of its upcoming models may reach the Critical threshold for cybersecurity under its Preparedness Framework. Preliminary internal evaluations of Astra showed enough progress that the company could not rule out that level.

The assessment marks a step up from earlier OpenAI systems, which were classified at the High threshold for frontier cyber risks rather than Critical. Astra is an unreleased model, and OpenAI said it was not involved in the exploitation of Hugging Face.

The judgement followed recent internal testing and expert assessments that pointed to advances in agentic coding and cybersecurity. Those results prompted OpenAI to increase robustness testing of its safeguards and security controls, and to tighten internal handling of the model.

Under OpenAI's framework, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or if it can devise and execute end-to-end novel strategies for cyberattacks against hardened targets from only a high-level goal.

OpenAI said Astra had not yet been definitively classified at that level. Instead, early evaluations were strong enough that the company could no longer exclude the possibility, prompting additional restrictions while benchmarking and assessment continue.

Controls tightened

The measures include stricter security controls for higher-capability models and related work. They cover isolated testing environments, restricted network and tool access, stronger protections and encryption for model weights, extra monitoring and detection systems, and sandboxed execution.

OpenAI has also paused internal activities involving Astra that do not meet the stronger security requirements. In addition, it has introduced universal monitoring for risky actions and misalignment across all agentic applications of the model, including during training and evaluation.

According to OpenAI, those monitors evaluate the model's chain of thought and can trigger a security response to review and interrupt high-risk activity. The company also said it would work with relevant government agencies and selected AI safety organisations to test the model, and would provide recommended security controls to third-party testing partners carrying out higher-risk evaluations and workloads.

Framework test

The disclosure offers a rare view into how one of the most closely watched AI developers is applying its preparedness process to cyber risk. OpenAI first published the Preparedness Framework in late 2023 as an internal guide for identifying advances in frontier model behaviour and deciding what safeguards and operational steps should follow.

The framework covers several areas of concern, including biology, chemistry, cybersecurity and AI self-improvement. OpenAI said previous transitions had already led it to strengthen safeguards in other domains, including biology, when its systems approached the high-capability threshold.

The cyber category is especially sensitive because more capable models can support both defence and offence. Security researchers and governments have increasingly focused on whether advanced AI tools might help identify software flaws, automate exploit development or co-ordinate attacks against hardened targets at a speed beyond current human teams.

At the same time, AI companies and cyber defenders argue that the same tools can be used to find and fix weaknesses before they are exploited. OpenAI framed its response in those terms, saying it believed advanced cyber-capable models should help defenders identify and address vulnerabilities before attackers do.

In its statement, the company said: "Cybersecurity is rapidly changing as models become more capable in ways that can both strengthen cyberdefenses and enable attacks at unprecedented speed and scale."

OpenAI said transparency was one reason for making the disclosure while testing remained under way. "We are sharing this because we believe it's important to be transparent with the public and the safety and security communities about this potential shift in capabilities," the company said.

It also set the current finding in the context of earlier model assessments. "Previous models, including GPT‐5.6‐Sol, have been evaluated for frontier cyber capabilities and assessed at the High (rather than Critical) threshold," OpenAI said.

On Astra's current status, OpenAI said: "While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time."

It added that development work would continue under tighter safeguards. "We are pausing internal activities involving Astra that do not yet meet these strengthened security control requirements," the company said.