MON, AUGUST 10, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ General

FLI AI Safety Index Summer 2026: No Lab Makes Honor Roll — Anthropic C+, xAI/DeepSeek/Mistral Failing

FLI Summer 2026 (July 7): Anthropic C+ (2.66/4.0) leads 9 labs across 6 domains. OpenAI C (2.28). Google DeepMind C (2.01). Meta D+. xAI, DeepSeek, Mistral F — one from each of US, China, Europe. No lab above C+. Key finding: Anthropic, OpenAI, Google DeepMind, and Meta all weakened or voided prior pause pledges — reviewers call it "moving goalposts." Existential safety: no company above C- in any lab.

By AIToolsRecap August 10, 2026 6 min read 307 views
Home Articles General Claude AI FLI AI Safety Index Summer 2026: Anthropic C+, ...

FLI AI SAFETY INDEX SUMMER 2026 — FULL RESULTS

Published: July 7, 2026 by the Future of Life Institute (FLI)
Coverage: 9 companies, 37 indicators, 6 domains, 7 independent expert reviewers
Anthropic: C+ (2.66/4.0) — #1 overall, leads 5 of 6 domains
OpenAI: C (2.28) — #2, leads Risk Assessment domain
Google DeepMind: C (2.01) — #3
Meta: D+ (1.32) — #4, improved from #6 last edition
Z.ai / Alibaba Cloud: D- — #5/#6
xAI / DeepSeek / Mistral: F — failing, one each from US/China/Europe
Existential safety: No company above C- — worst domain across all labs
Key finding: Top 4 labs weakened or voided pause pledges — reviewers call it "moving goalposts"
The gap: Anthropic (2.66) vs OpenAI (2.28) vs GDM (2.01) are close. GDM to Meta (1.32) is a 0.7-point drop. Failing labs cluster at 0.33-0.65.

What the FLI Safety Index Measures — and What It Does Not

Per the official FLI report, the index grades labs across six domains: Risk Assessment, Current Harms, Safety Frameworks, Existential Safety, Governance and Accountability, and Information Disclosure. Critically, as Digital Applied's enterprise buyer analysis clarifies, the index "grades labs, not deployments" — it evaluates policies, governance structures, and published frameworks, not the performance of specific deployed products or user experiences. A C+ does not mean Claude is safer to use than GPT-5.6 Sol on any given task. It means Anthropic's institutional posture on safety — its governance, disclosures, and commitments — is better documented and more robust than its peers. For enterprise procurement decisions, this distinction matters.

The Most Significant Finding — Moving the Goalposts

According to TechTimes's analysis of the FLI report, the most consequential finding is not the grades themselves but what the top-ranked companies have stopped promising. Anthropic, OpenAI, Google DeepMind, and Meta — all of which previously had explicit commitments to pause development if dangerous capability thresholds were approached — have weakened or voided those pledges. As the FLI panel notes, some labs now use "competitor-contingent language" — essentially committing to pause only if competitors also pause. Reviewers called this "moving the goalposts" and argued that it "undermines safety frameworks' credibility." This finding lands one week after OpenAI paused Astra because internal evaluations triggered its Critical cybersecurity threshold — exactly the kind of situation those pause pledges were designed for. OpenAI acted on its Preparedness Framework in the Astra case. The FLI finding is that the broader pause pledge architecture is weakening across the industry even as capability thresholds are being reached.

Full Scorecard

LabGradeScore (4.0 scale)Notes
AnthropicC+2.66Leads 5 of 6 domains. Criticised for military engagements.
OpenAIC2.28Leads Risk Assessment. Slipped from C+ last edition.
Google DeepMindC2.01#3 overall.
MetaD+1.32Improved from 6th to 4th place this edition.
Z.ai / Alibaba CloudD-0.65 / 0.33Deny US military ties allegations.
xAI · DeepSeek · MistralFFailingOne from US, China, Europe. xAI dropped from 4th to 7th.

What It Means for Enterprise Buyers

As Digital Applied's enterprise readout advises, the FLI grades are useful as a policy and governance signal but should not be the sole input to a procurement decision. The index explicitly evaluates institutional behaviour, not deployed product safety. "A C+ says nothing about whether your specific model, configuration, and contract meet your bar." For enterprise AI procurement, the practical takeaway is: Anthropic leads on documented safety governance. OpenAI and Google DeepMind are close behind on the same scale. xAI, DeepSeek, and Mistral have failed to demonstrate adequate safety frameworks by the panel's standards. For regulated industries where vendor safety posture is a procurement requirement, the FLI index provides the clearest third-party signal available — use it alongside your own audit.

The index also lands with the Astra pause still fresh. OpenAI paused Astra because its Preparedness Framework's Critical threshold was triggered. The FLI finding says the broader industry is weakening those same threshold frameworks. These two data points in the same week — one lab acting on its framework, multiple labs softening theirs — represent the clearest picture yet of where the AI safety governance gap actually sits.

Sources: FLI official report (July 7, 2026) · AI Weekly full coverage · Digital Applied enterprise analysis · ThePlanetTools detailed grades · BigGo Finance · Related: OpenAI pauses Astra: Critical cyber threshold → · Anthropic CGAO hire →

Tags
AI NewsGenerative AIAnthropicClaude Code2026

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →