FLI AI SAFETY INDEX SUMMER 2026 — FULL RESULTS
● Published: July 7, 2026 by the Future of Life Institute (FLI)
● Coverage: 9 companies, 37 indicators, 6 domains, 7 independent expert reviewers
● Anthropic: C+ (2.66/4.0) — #1 overall, leads 5 of 6 domains
● OpenAI: C (2.28) — #2, leads Risk Assessment domain
● Google DeepMind: C (2.01) — #3
● Meta: D+ (1.32) — #4, improved from #6 last edition
● Z.ai / Alibaba Cloud: D- — #5/#6
● xAI / DeepSeek / Mistral: F — failing, one each from US/China/Europe
● Existential safety: No company above C- — worst domain across all labs
● Key finding: Top 4 labs weakened or voided pause pledges — reviewers call it "moving goalposts"
● The gap: Anthropic (2.66) vs OpenAI (2.28) vs GDM (2.01) are close. GDM to Meta (1.32) is a 0.7-point drop. Failing labs cluster at 0.33-0.65.
What the FLI Safety Index Measures — and What It Does Not
Per the official FLI report, the index grades labs across six domains: Risk Assessment, Current Harms, Safety Frameworks, Existential Safety, Governance and Accountability, and Information Disclosure. Critically, as Digital Applied's enterprise buyer analysis clarifies, the index "grades labs, not deployments" — it evaluates policies, governance structures, and published frameworks, not the performance of specific deployed products or user experiences. A C+ does not mean Claude is safer to use than GPT-5.6 Sol on any given task. It means Anthropic's institutional posture on safety — its governance, disclosures, and commitments — is better documented and more robust than its peers. For enterprise procurement decisions, this distinction matters.
The Most Significant Finding — Moving the Goalposts
According to TechTimes's analysis of the FLI report, the most consequential finding is not the grades themselves but what the top-ranked companies have stopped promising. Anthropic, OpenAI, Google DeepMind, and Meta — all of which previously had explicit commitments to pause development if dangerous capability thresholds were approached — have weakened or voided those pledges. As the FLI panel notes, some labs now use "competitor-contingent language" — essentially committing to pause only if competitors also pause. Reviewers called this "moving the goalposts" and argued that it "undermines safety frameworks' credibility." This finding lands one week after OpenAI paused Astra because internal evaluations triggered its Critical cybersecurity threshold — exactly the kind of situation those pause pledges were designed for. OpenAI acted on its Preparedness Framework in the Astra case. The FLI finding is that the broader pause pledge architecture is weakening across the industry even as capability thresholds are being reached.
Full Scorecard
| Lab | Grade | Score (4.0 scale) | Notes |
| Anthropic | C+ | 2.66 | Leads 5 of 6 domains. Criticised for military engagements. |
| OpenAI | C | 2.28 | Leads Risk Assessment. Slipped from C+ last edition. |
| Google DeepMind | C | 2.01 | #3 overall. |
| Meta | D+ | 1.32 | Improved from 6th to 4th place this edition. |
| Z.ai / Alibaba Cloud | D- | 0.65 / 0.33 | Deny US military ties allegations. |
| xAI · DeepSeek · Mistral | F | Failing | One from US, China, Europe. xAI dropped from 4th to 7th. |
What It Means for Enterprise Buyers
As Digital Applied's enterprise readout advises, the FLI grades are useful as a policy and governance signal but should not be the sole input to a procurement decision. The index explicitly evaluates institutional behaviour, not deployed product safety. "A C+ says nothing about whether your specific model, configuration, and contract meet your bar." For enterprise AI procurement, the practical takeaway is: Anthropic leads on documented safety governance. OpenAI and Google DeepMind are close behind on the same scale. xAI, DeepSeek, and Mistral have failed to demonstrate adequate safety frameworks by the panel's standards. For regulated industries where vendor safety posture is a procurement requirement, the FLI index provides the clearest third-party signal available — use it alongside your own audit.
The index also lands with the Astra pause still fresh. OpenAI paused Astra because its Preparedness Framework's Critical threshold was triggered. The FLI finding says the broader industry is weakening those same threshold frameworks. These two data points in the same week — one lab acting on its framework, multiple labs softening theirs — represent the clearest picture yet of where the AI safety governance gap actually sits.
Sources: FLI official report (July 7, 2026) · AI Weekly full coverage · Digital Applied enterprise analysis · ThePlanetTools detailed grades · BigGo Finance · Related: OpenAI pauses Astra: Critical cyber threshold → · Anthropic CGAO hire →