OpenAI's Astra Becomes First Model to Hit Critical Cybersecurity Threshold — With Stronger Safeguards

OpenAI details safety work on Astra, the first model designated at Critical cybersecurity capability under its Preparedness Framework, achieving 100% on ExploitBench.

Tuesday September 1, 2026 Source: OpenAI
TL;DR — Quick Answer

OpenAI detailed safety work on Astra, its first model to hit the Critical cybersecurity threshold under its Preparedness Framework. Astra scored 100% on ExploitBench, discovered two real zero-days during testing, and built full browser and OS exploit chains — prompting delayed development, 91.5% jailbreak refusal, and restricted access for advanced cyber capabilities.

Key Takeaways

OpenAI's Astra Becomes First Model to Hit Critical Cybersecurity Threshold — With Stronger Safeguards — AI news article illustration

OpenAI has published a detailed safety assessment of Astra, its upcoming frontier model — and the first model the company has designated at the Critical cybersecurity capability threshold under its Preparedness Framework. The announcement means Astra can find previously unknown security flaws and develop exploits across well-protected systems without a person guiding each step.

What Critical Capability Means

Under OpenAI’s Preparedness Framework, a model meets the Critical threshold if it can:

Astra is the first OpenAI model to reach this designation, requiring stronger safeguards during development and before release.

Benchmark Results

The evaluation results are striking:

Safeguards for Release

OpenAI delayed parts of Astra’s development while strengthening protections:

Restricted Access

Access to Astra’s most advanced cybersecurity capabilities will be limited:

What It Means

The announcement marks a significant moment for AI safety: the first model where cybersecurity capability itself is the frontier concern. OpenAI’s transparency about the risks — including the possibility of slowed legitimate work — reflects the seriousness of deploying a model that can autonomously find and exploit zero-day vulnerabilities.

Astra will be available soon, with full details in the model’s system card at launch.

Frequently Asked Questions

What is OpenAI's Critical cybersecurity threshold?

Under OpenAI's Preparedness Framework, a model meets the Critical threshold if it can develop functional zero-day exploits in hardened real-world systems without human intervention, or execute end-to-end novel cyberattack strategies from just a high-level goal. Astra is the first model designated at this level.

What did Astra score on cybersecurity benchmarks?

Astra achieved a perfect 100% on ExploitBench for developing exploits from known vulnerabilities, and on an internal benchmark of 20 recent high-severity V8 vulnerabilities it achieved far higher code-execution rates than GPT-5.6 Sol with far fewer tokens. During evaluation it even discovered and used two real zero-day vulnerabilities.

Is Astra dangerous?

OpenAI's assessment is that Astra's safeguards sufficiently minimize severe-harm risk for release, but the model's capability level required stronger protections: development delays, 91.5% jailbreak refusal rates, chain-of-thought misalignment monitoring in production, and restricted access to its most advanced cyber features through the Daybreak program.

When will Astra be available?

OpenAI stated Astra will be available soon, with advanced cybersecurity capabilities limited initially to a small group of alpha testers, followed by broader defensive access through Daybreak Blue. Full details will be in the model's system card at launch.

This article is based on the official announcement from OpenAI . Read the original for full technical details.

Related Articles

Back to all news