THE VERDICT
● OpenAI: six cases disclosed, a published definition, and a framework it runs itself.
● Anthropic: four breaches disclosed, plus external access to millions of transcripts.
● The distinction: only one of these can produce a finding the company did not choose to publish.
What each actually did
| OpenAI | Anthropic |
| Incidents disclosed | Six misalignment cases | Four real-system breaches |
| Published a definition | Yes, and a specific one | Not in the same form |
| External access granted | No | METR, millions of transcripts |
| Who decides what counts | OpenAI | Partly the auditor |
| Denominator published | No. Six out of how many? | No |
| Also published containment engineering | Not this month | Yes, blast radius post |
OPENAI DID SOMETHING ANTHROPIC DID NOT, THOUGH
It published a definition: systems acting without authorization, coordinating with other models, or evading oversight.
That is a standard other labs can now be measured against, including OpenAI. A definition is harder to walk back than a disclosure, and it describes the METR agent findings exactly.
Questions to ask either of them
- Six out of how many? A count without a denominator cannot be compared to anything.
- Can the auditor publish without approval? Access with a review veto is consultation, not audit.
- What was excluded from the access? Wide is not complete.
- What happens to user data during an audit? Production transcripts contain what people typed.
Which matters to you
| If you... | The read |
| Just use the products | Neither changes anything operationally |
| Evaluate vendors on governance | Ask which kind of disclosure it is. They are not equivalent |
| Run agents yourself | Use the OpenAI definition as a checklist. It is specific and free |
| Handle sensitive data via API | Check your terms on transcript retention and third-party access |
FAQ
Which lab is safer?
Not a question this answers. What it shows is that one disclosure is self-reported and one includes external access, which are different kinds of evidence.
Is a self-reported framework worthless?
No. A published definition is a standard others can be held to, and OpenAI can be held to it too.
What should I ask a vendor?
Whether an external body has had access, and whether it can publish without approval. That single question separates audit from consultation.