What OpenAI Published
OpenAI states that Astra is the first of its models to meet the Critical cybersecurity capability threshold under its Preparedness Framework.
That threshold is defined by the ability to do one of two things: identify and develop working zero-day exploits against hardened systems without human intervention, or devise and execute novel attack strategies against hardened targets.
The Numbers
- 100% on ExploitBench. A perfect score on the public benchmark.
- Two zero-days found during evaluation. Astra discovered and then used two previously unknown vulnerabilities as part of exploit chains.
- 91.5% refusal rate on cyber jailbreak evaluations.
- Higher code-execution rates than GPT-5.6 Sol using far fewer output tokens, on an internal port of ExploitBench run between June and August 2026.
OpenAI's own comparison baseline is GPT-5.6 Sol. That is the figure they published, and it is the one used here.
Reading the 91.5% Honestly
A 91.5% refusal rate is the headline safety number, and it is worth being clear about what it means.
It means roughly one in twelve cyber jailbreak attempts got through in evaluation. For a model that scores 100% on exploit generation, that residual is the entire risk surface. Whether it is acceptable depends completely on who holds access - which is why the access model below matters more than the benchmark.
The Safeguards
OpenAI describes a layered approach rather than a single control:
- Training the model to refuse harmful cyber requests more reliably
- System-level classifiers that detect cyber abuse
- Chain-of-thought monitoring to catch misaligned actions mid-task
- More conservative model boundaries for higher-risk accounts
The last one is the significant one. It means the same model behaves differently depending on who is asking - a policy decision implemented in the product rather than a property of the weights.
Who Gets Access
Astra starts with alpha testers, then widens through a programme called Daybreak Blue aimed at defensive use. OpenAI says it plans to make Astra available soon without giving a date.
The logic is that defenders benefit more than attackers from a tool that finds vulnerabilities, because defenders can fix what they find and attackers already have working methods. That argument is genuinely contested among security researchers, and it is the argument the whole release strategy rests on.
The Industry Moved Together This Month
Astra did not arrive alone. Google announced Gemini 3.8 Flash Cyber through a programme called Fairwind, giving early access to high-priority defenders such as governments, healthcare providers and telecoms, working with more than 650 partners. Anthropic introduced Enterprise Frontier Safeguards pairing zero data retention with misuse detection, and now permits Claude Fable 5.1 for vulnerability identification while routing penetration testing and exploit generation to other models.
Three labs published cyber capability disclosures and access-control programmes within weeks of each other, and days before the White House AI meeting on 29 September. Read that sequence as the labs demonstrating they can govern this themselves.
What It Means Practically
For most builders: nothing yet, because you cannot get access. For security teams, the thing to plan for is not Astra itself but what it implies - autonomous vulnerability discovery is now a demonstrated capability, and the gap between a model that can find a zero-day and one that is widely available is an access policy, not a technical barrier.
Patch cadence assumptions built on the idea that finding novel vulnerabilities is slow and expensive need revisiting.
FAQ
What is the Critical threshold?
The highest cybersecurity tier in OpenAI's Preparedness Framework, met when a model can autonomously develop working zero-day exploits against hardened systems, or devise and execute novel attacks against hardened targets.
Can I use Astra?
Not yet. Access begins with alpha testers and expands through the Daybreak Blue defensive-use programme. No public date has been given.
What is ExploitBench?
A benchmark measuring exploit development capability. Astra scored 100%. OpenAI also ran an internal port between June and August 2026.
Did Astra really find zero-days?
OpenAI states it discovered and used two previously unknown vulnerabilities in exploit chains during evaluation.