TUE, SEPTEMBER 29, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ General

Astra Found Two Zero-Days, Then OpenAI Called It Critical

OpenAI says Astra is the first model to cross the Critical cybersecurity threshold in its Preparedness Framework, scoring 100% on ExploitBench and refusing 91.5% of cyber jailbreak attempts, with access limited to alpha testers and a defensive-use programme.

By AIToolsRecap September 29, 2026 6 min read 16 views
Home › Articles › General › OpenAI Astra Is the First Model to Hit Critical

What OpenAI Published

OpenAI states that Astra is the first of its models to meet the Critical cybersecurity capability threshold under its Preparedness Framework.

That threshold is defined by the ability to do one of two things: identify and develop working zero-day exploits against hardened systems without human intervention, or devise and execute novel attack strategies against hardened targets.

The Numbers

  • 100% on ExploitBench. A perfect score on the public benchmark.
  • Two zero-days found during evaluation. Astra discovered and then used two previously unknown vulnerabilities as part of exploit chains.
  • 91.5% refusal rate on cyber jailbreak evaluations.
  • Higher code-execution rates than GPT-5.6 Sol using far fewer output tokens, on an internal port of ExploitBench run between June and August 2026.

OpenAI's own comparison baseline is GPT-5.6 Sol. That is the figure they published, and it is the one used here.

Reading the 91.5% Honestly

A 91.5% refusal rate is the headline safety number, and it is worth being clear about what it means.

It means roughly one in twelve cyber jailbreak attempts got through in evaluation. For a model that scores 100% on exploit generation, that residual is the entire risk surface. Whether it is acceptable depends completely on who holds access - which is why the access model below matters more than the benchmark.

The Safeguards

OpenAI describes a layered approach rather than a single control:

  • Training the model to refuse harmful cyber requests more reliably
  • System-level classifiers that detect cyber abuse
  • Chain-of-thought monitoring to catch misaligned actions mid-task
  • More conservative model boundaries for higher-risk accounts

The last one is the significant one. It means the same model behaves differently depending on who is asking - a policy decision implemented in the product rather than a property of the weights.

Who Gets Access

Astra starts with alpha testers, then widens through a programme called Daybreak Blue aimed at defensive use. OpenAI says it plans to make Astra available soon without giving a date.

The logic is that defenders benefit more than attackers from a tool that finds vulnerabilities, because defenders can fix what they find and attackers already have working methods. That argument is genuinely contested among security researchers, and it is the argument the whole release strategy rests on.

The Industry Moved Together This Month

Astra did not arrive alone. Google announced Gemini 3.8 Flash Cyber through a programme called Fairwind, giving early access to high-priority defenders such as governments, healthcare providers and telecoms, working with more than 650 partners. Anthropic introduced Enterprise Frontier Safeguards pairing zero data retention with misuse detection, and now permits Claude Fable 5.1 for vulnerability identification while routing penetration testing and exploit generation to other models.

Three labs published cyber capability disclosures and access-control programmes within weeks of each other, and days before the White House AI meeting on 29 September. Read that sequence as the labs demonstrating they can govern this themselves.

What It Means Practically

For most builders: nothing yet, because you cannot get access. For security teams, the thing to plan for is not Astra itself but what it implies - autonomous vulnerability discovery is now a demonstrated capability, and the gap between a model that can find a zero-day and one that is widely available is an access policy, not a technical barrier.

Patch cadence assumptions built on the idea that finding novel vulnerabilities is slow and expensive need revisiting.

FAQ

What is the Critical threshold?

The highest cybersecurity tier in OpenAI's Preparedness Framework, met when a model can autonomously develop working zero-day exploits against hardened systems, or devise and execute novel attacks against hardened targets.

Can I use Astra?

Not yet. Access begins with alpha testers and expands through the Daybreak Blue defensive-use programme. No public date has been given.

What is ExploitBench?

A benchmark measuring exploit development capability. Astra scored 100%. OpenAI also ran an internal port between June and August 2026.

Did Astra really find zero-days?

OpenAI states it discovered and used two previously unknown vulnerabilities in exploit chains during evaluation.

Tags
AI NewsOpenAI2026
⚑

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →
💡 AI Tools prompts
Prompt Guide
Best Claude AI Prompts for SEO (2026) — Content, Technical, and Comparison SEO
Claude Sonnet 5 and Opus 5 are strong for SEO work that requires writing quality, structured analysis, and long-form content generation. With 1M context, Claude can analyse an entire site's content structure, compare competing pages, and write complete article drafts in one session. These prompts cover the full SEO workflow: keyword research synthesis, content briefs, on-page optimisation, meta descriptions, technical audit interpretation, and comparison content that ranks above AI Overviews.
Get Prompts →
Prompt Guide
Best ChatGPT Prompts for SEO (2026) — GPT-5.6 and Browse
ChatGPT with GPT-5.6 Sol and Browse enabled is a capable SEO research tool — it can search the live web, analyse SERP results, and synthesise content briefs in a single session. GPT-5.6 Terra at $2.50/M offers a cost-efficient option for high-volume SEO content generation. These prompts are optimised for ChatGPT Plus with Browse, the ChatGPT Work product for larger projects, and the OpenAI API with web_search tool enabled.
Get Prompts →
Prompt Guide
Best Claude Opus 5 and Sonnet 5 Prompts for Writing (2026)
Claude Opus 5 and Sonnet 5 consistently produce the highest-quality long-form writing of any AI model in July 2026 — a lead documented across writing benchmarks and user testing since Claude 3 Opus. With 1M context and 128K output on Opus 5, Claude can write book chapters, complete reports, and long-form content without truncating. Sonnet 5 at $2/$10/M (intro through August 31) is the best value writing model available. These prompts are optimised for claude.ai Pro/Max, Claude Cowork, and the API.
Get Prompts →