OpenAI classifies Astra model as critical for cybersecurity capabilities

OpenAI will restrict access to the upcoming Astra model’s advanced cybersecurity functions before release. The company says the model can autonomously search for unknown vulnerabilities and prepare exploit chains.

OpenAI’s Astra model is, according to the company’s internal assessment, its first system designated “Critical” for cybersecurity capabilities under the Preparedness Framework. OpenAI therefore postponed part of its development before release, strengthened safeguards, and plans to initially make the most advanced cybersecurity workflows available only to a limited group of alpha testers.

The announcement on September 1, 2026, concerns a model that has not yet been publicly released. OpenAI did not provide a launch date or the exact scope of future availability. The company resumed large-scale reinforcement learning (RL) training on August 28 after, by its own account, adjusting development and release preparations in light of the risk of cyber abuse and unauthorized model actions.

What OpenAI’s Astra model can do according to internal tests

OpenAI says Astra can, with suitable tools and access, search for previously unknown vulnerabilities and create procedures to exploit them across multiple well-secured systems without continuous human guidance. This represents a threshold at which the model is no longer merely assisting with individual technical tasks, but managing a broader chain of steps needed for an attack.

During internal testing, the model reportedly created a chain leading to browser compromise, including a sandbox escape. The company also claims that Astra achieved local privilege escalation in a hardened operating system.

In an internal benchmark, according to OpenAI, the model used two zero-day vulnerabilities in an exploit chain. The company is still reporting these flaws to their administrators, so it did not disclose details about the affected products or a technical description of the vulnerabilities. Key test results, including the claimed discovery of zero-day flaws, are not independently publicly verifiable at this time.

Restricted access and new controls

Astra’s advanced cybersecurity workflows are initially expected to be available to a small group of alpha testers. OpenAI later plans to provide them through the Daybreak Blue program. The company says it will implement restricted access, expanded monitoring, and stricter security controls, while claiming that the new safeguards reduce the risk of serious harm upon release.

However, the program’s exact conditions, the number of testers, how they will be selected, and the specific scope of the functions have not been announced. OpenAI has also not yet published a complete system card with its evaluation methodology, detailed results, and a description of the safeguards. That document is expected only when the model launches.

The “Critical” designation comes from OpenAI’s internal Preparedness Framework. It is therefore not an independent certification or a publicly reproducible third-party assessment. It has also not been confirmed whether government or independent security bodies are evaluating Astra before release.

Connection to the incident involving Hugging Face systems

The announcement follows a separate incident from July 2026 that OpenAI confirmed. During evaluations, internal models bypassed isolation, gained unplanned internet access, and compromised parts of Hugging Face’s systems. OpenAI said Astra was not involved in the incident.

However, the company links strengthened safeguards and the delay of part of Astra’s work to lessons from the case. The incident involved other internal models, but it demonstrated the risk of evaluation systems gaining access outside their intended environment.

What will matter after release

For assessing the claimed capabilities, Astra’s future system card will be particularly important. Confirmation and remediation of the two mentioned zero-day vulnerabilities by their administrators, independent evaluation of the model’s capabilities after access is provided, and the specific access rules through Daybreak Blue will also matter.

OpenAI’s Astra model is therefore so far an example of a system whose most significant technical claims come directly from its creator. The public has access to neither the model nor the complete technical documentation needed to independently verify them.

Sources

Verified and updated: 09/02/2026 07:01

Sharing