SAT, SEPTEMBER 12, 2026
Independent · In‑Depth · Practitioner‑Tested
General

How to Tell a Real AI Safety Commitment From a Press Release

Anthropic gave METR its transcripts. OpenAI gated a capability behind an application. Most labs publish a document. Those are not the same thing.

🕐 5 min read 👁 36 views 📅 Sep 12, 2026
THE VERDICT

● The test: could this commitment produce a finding the company did not want?

● If yes, it costs something and it means something.

● If no, it is a description of work already done, graded by the people who did it.

Four kinds, ranked by what they cost

KindExampleCan it produce unwanted findings?
External audit with accessAnthropic giving METR transcriptsYes, if publication is independent
Disclosure against interestAnthropic publishing four breachesAlready did
Capability gatingOpenAI routing Astra cyber via DaybreakCosts revenue. Enforcement unpublished
Framework documentMost published safety policiesRarely
THE DETAIL THAT DECIDES THE TOP ROW

External access only counts if the auditor publishes independently. Access with a review veto is consultation wearing an audit costume.

That detail has not been published for the METR arrangement, and it is the first thing to ask about any similar commitment.

Questions to ask any vendor

  • Has an external body had access? To what, and could they publish without approval?
  • Have you disclosed anything against your own interest? If not, the safety documentation is marketing.
  • What have you declined to ship? A lab that has never withheld anything has never had to choose.
  • How is any gating enforced technically? Policy and control are different, and only one survives an incident.
  • What happens to my data in an audit? Production transcripts contain what your users typed.

FAQ

Is a published safety framework worthless?

No, but it is the weakest of the four. It describes intent rather than demonstrating a cost, and every lab has one.

What makes external audit different?

It can produce findings the company did not want, provided the auditor publishes independently. Without that, it is consultation.

Which labs have done the strongest version?

Anthropic has both disclosed breaches against its own interest and granted external transcript access. OpenAI has gated a capability at revenue cost. Both are more than a document, and neither has published full enforcement detail.

⚖ Our Verdict

The test is whether a commitment could produce a finding the lab did not want. Most cannot.