TUE, AUGUST 18, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ News

The Filter Was Off for 11 Months. So Was the Alarm That Would Have Told Anyone

Anthropic's blocking biological classifiers did not run on any human-feedback platform traffic from May 2025 to April 2026, covering roughly 133 million exchanges from about 50,000 contractors vetted by outside vendors. An internal-only flag disabled the classifiers and their logging together, so missed blocks left no record. A retroactive sweep flagged 1,197 transcripts and found no clearly concerning misuse.

By AIToolsRecap August 17, 2026 9 min read 12 views
Home Articles News Claude AI Anthropic Ran 133 Million Contractor Chats With...
THE FACTS

● Duration: May 2025 to April 2026. Eleven months.

● Scope: roughly 133 million exchanges from about 50,000 human-feedback contractors. Customer traffic was not affected.

● The compounding failure: an internal-only flag disabled the blocking and the logging together. Missed blocks left no record.

● Retroactive review: 1,197 transcripts flagged high risk. Manual review found no clearly concerning misuse.

● Consequence: Anthropic downgraded its own February safety verdict from very low to low, retroactively.

What the safeguard was supposed to do

Anthropic runs blocking classifiers intended to stop its models assisting anyone with chemical or biological weapons work. It first deployed models carrying those safeguards in May 2025.

From that point until April 2026, the classifiers did not run on any traffic passing through the company's human feedback platforms — the systems where external contractors interact with models to generate training signal.

The numbers

Metric Figure
Duration 11 months, May 2025 to April 2026
Exchanges ~133 million
People with access ~50,000 contractors
Who vetted them Outside vendors, not Anthropic
Interaction type Vast majority could hold open-ended conversations, not just rate fixed answers
Customers affected None, per the report

On the vetting, the report is blunt: many of those vendors did not have screening processes capable of stopping even the lowest-tier threat actors in Anthropic's own classification scheme.

The part that turns a gap into a blind spot

THE LOGGING WENT DOWN WITH IT

This was not a jailbreak slipping past a working safeguard. The safeguard was not running on those platforms — and neither was its alert trail. An internal-use flag disabled both the real-time classifiers and the logging of their flags simultaneously. A missed block would not have left a record for anyone to review later.

That is the difference between a control that failed and a control whose failure was invisible. The second kind is what lets eleven months pass.

What the retroactive review found

After discovering the gap, Anthropic ran Claude Sonnet 5 over every human turn sent during the affected period. Results:

  • 1,197 transcripts flagged as high risk
  • 757 of those came from Anthropic's own teams
  • Of the remainder, all but 62 came from deliberate red-teaming
  • Manual review covered those 62 plus 30 random red-teaming transcripts
  • Finding: no clearly concerning misuse, though some dual-use conversations were flagged

Anthropic's own conclusion is the notable part. Having found this, the company now says it believes similar undiscovered issues are more likely than it previously assumed. That is a lab reasoning correctly about what one discovered failure implies about the base rate of undiscovered ones.

The second incident on the same platforms

An outside tip arrived in April 2026. Anthropic confirmed that a small number of contractors at data-labelling vendors had exploited a flaw to obtain an API key, then used models outside their assigned work. One of those was Mythos Preview, among Anthropic's most capable systems, running without blocking biological classifiers for roughly two of the several weeks the access path stayed open.

Containment was fast once known: within 90 minutes of learning of it, with the vector closed the same day. No model weights were taken, no customer data was reached.

The pattern across both incidents is the same, and it is not about model capability. It is about the vendor pathway — identity checks, credential issuance, classifier routing, logging, alerting — working as one system rather than five independently plausible ones.

Anthropic corrected its own previous report

The August report retroactively downgrades the February verdict. Risk from non-novel weapons uplift is now assessed as low but higher than the previous estimate, where February had said very low.

A safety report that corrects the preceding safety report is genuinely uncommon. It is also the strongest available evidence that the series is not a marketing exercise, since a document designed to reassure would not open by revising its predecessor downward.

What to actually do with this

If you are... Do this
An Anthropic API customer Nothing. Customer traffic was outside the affected pathway
Running your own safety classifiers Alert on classifier silence, not just on classifier hits. Zero flags is an alarm state
Using outsourced labelling or feedback vendors Audit who does the vetting and against what standard. Outsourced vetting is your risk, not theirs
Writing AI vendor due diligence Ask whether internal feature flags can disable a control and its logging together
Comparing labs on safety Ask what the others have not published. Silence is not a clean record

The disclosure incentive problem

Anthropic looks worse today than every lab that publishes nothing comparable. That is a bad structure. The company that measures itself hardest and publishes the results takes the reputational hit, while the ones producing no equivalent document take none.

The check that could fix this is external. Anthropic's Long-Term Benefit Trust holds the power to compel independent review of future risk reports. Whether it exercises that power is the thing to watch, because self-disclosure without external verification only works as long as the disclosing party keeps choosing to disclose.

FAQ

Was customer data or customer traffic affected?

No. The gap covered human-feedback contractor platforms, not customer API or product traffic. In the separate API key incident, no model weights were taken and no customer data was reached.

Did anyone actually misuse the gap?

The retroactive review found no clearly concerning misuse. Of 1,197 high-risk flags, 757 were Anthropic's own teams and nearly all the rest were deliberate red-teaming. Manual review of the remaining 62 plus a random sample turned up dual-use conversations but nothing clearly concerning.

How did this go unnoticed for eleven months?

Because the flag that disabled the classifiers also disabled the logging of their flags. There was no signal to notice. This is the operational lesson: a monitoring system that can be silenced by the same switch that silences the control has no independent failure mode.

What is Mythos Preview?

An unreleased, highly capable Anthropic model. Contractors who obtained an API key through the exploited flaw reached it for roughly two weeks, during which it ran without blocking biological classifiers.

Has it been fixed?

Per the report, yes. The gap has been remediated and contractor requirements tightened. Anthropic also says it now considers similar undiscovered issues more likely than it previously believed.

Does this mean Anthropic is less safe than other labs?

It means Anthropic published a 186-page self-assessment including its own failures. No comparable document exists from most of its peers, so the honest answer is that there is nothing to compare it against.

Tags
AnthropicClaudeAI SafetyMythos PreviewClaude Sonnet 5Responsible Scaling PolicySecurityAI governance2026

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →