OpenAI has published the full capabilities announcement for GPT-6 Astra, calling it “a new generation of intelligence” — the world’s most intelligent and aligned model, state-of-the-art across computer use, browsing, software engineering, cybersecurity, science, and professional work.
Saturated Benchmarks, New Mathematics
The headline numbers reframe what frontier models can do:
- FrontierMath Tier 4: 98% — and Astra has “already helped solve long-standing open problems in mathematics”
- ARC-AGI-3: 99.9% — surpassing the human action-efficiency baseline on 96% of levels, which the ARC Prize Foundation called “a meaningful step change in frontier-model performance”
- ExploitBench: 100% — a perfect score turning known vulnerabilities into working exploits (vs 78.5% for GPT-5.6 Sol)
OpenAI shared two new results on prime gaps that Astra helped establish: a proof that infinitely many prime pairs occur within 186 of each other (improving the decade-old 246 bound), and an improved term in the large-gap bound that had stood for over 80 years. Both proofs and verification materials are published.
The World’s Best Computer Use
Astra sets a new frontier on computer and browser use: 72.6% on OSWorld 2.0 at roughly 40 minutes per task — versus GPT-5.6 Sol’s 65.7% at roughly 75 minutes, a 47% time reduction. On Agents’ Last Exam, it scores 59.3% using ~65% fewer output tokens than Claude Opus 5.
Demonstrated tasks include PCB layout in KiCad, Form 1040 completion, frontend QA, Power BI analysis, and Blender-to-Unreal Engine scene creation.
Alignment: Learning From Incidents
The post details a new evaluation built from the Hugging Face incident, testing whether a model facing an impossible task goes beyond its authorized scope. GPT-5.6 Sol (without production safeguards) did so 48% of the time. GPT-6 Astra: 0%.
Astra never attempted to circumvent Codex Auto-Review denials — even when auto-review was deliberately configured as evadable. It’s also 3x less likely to make inaccurate claims about its own capabilities.
OpenAI acknowledges Astra’s written reasoning is harder to monitor than Sol’s in adversarial tests, attributing this to fewer written reasoning steps, and states that improving monitorability remains a research priority.
Coding and Codex Upgrades
Astra is OpenAI’s best software engineering model: 57.9% on Terminal-Bench 4.0, 64.5% on FrontierCode Extended, and 63.9% on internal database migration tasks.
Codex gets a new context-preservation system: instead of lossy compaction summaries, Astra keeps notes across context windows and previous windows remain searchable — so it can retrieve requirements or test results from earlier in a long session.
Availability and Pricing
GPT-6 Astra is rolling out today to a limited set of organizations, then over coming days to all ChatGPT Plus, Pro, Business, and Enterprise users, the OpenAI API (as gpt-6-astra), Microsoft Azure, and AWS Bedrock.
- API pricing: $10/M input, $50/M output; Fast mode at 2x speed for 2x price
- GPT-6 Astra Pro for Pro, Business, and Enterprise plans
- Zero Data Retention supported for eligible API customers
- Enterprise workspaces: admin-enabled, off by default at launch
The launch escalates the frontier race dramatically — arriving just two days after Anthropic’s Fable 5.1 and Mythos 5.1, with benchmark comparisons showing Astra leading on most measures while Claude Fable 5.1 wins Humanity’s Last Exam (65.0% vs 57.2%).