THE 60-SECOND VERSION
● Anthropic raised its own misalignment rating from very low to low, and disclosed Model 2 — more capable than Mythos 5, with no plans to release it.
● The buried story: bioweapon classifiers were off for 11 months across 133 million contractor exchanges. Logging was off too.
● 81 percent of Americans aged 18 to 34 say they do not trust Alex Karp on AI. Satya Nadella was the only executive tested with net trust.
● 13 days left on Claude Sonnet 5 at 2 dollars per million input. September 1 it becomes 3 dollars.
Anthropic marked its own homework down
Anthropic published its second company-wide Risk Report on August 14, 2026, under version 3.4 of its Responsible Scaling Policy. It runs 186 pages and covers February 24 through a July 15 cutoff. It is the first to assess internal-only models alongside shipped ones.
The headline change moves in the wrong direction. Catastrophic harm from misalignment in high-stakes settings went from very low to low. The cause is worth stating precisely, because most coverage got it backwards: no model failed a safety test. Anthropic says recent cybersecurity-evaluation incident disclosures raised its overall uncertainty, and it moved the label even though it believes its underlying argument probably still supports very low.
MODEL 2
A Mythos-class model, more capable than Mythos 5 on Anthropic's own engineering benchmark, used heavily inside the company. The report states plainly: no current plans to release it externally. Its internal deployment review surfaced no new or more concerning misalignment than Mythos 5 already shows.
The item getting the least attention may matter most. Anthropic's task-based evaluations for automated AI research and development have saturated — they no longer register capability gains — at the same moment the company reports early signs of acceleration. The instrument built to detect a dangerous threshold crossing can no longer measure movement toward it. The rating stayed at low, but with less confidence than before.
Internally, Claude now writes a large majority of the code merged into Anthropic's production codebases, and the company estimates AI-assisted R&D is meaningfully faster than unaided work, though not yet twice as fast.
We broke down the full report separately, including the agent behaviours Anthropic classifies as apparent-success-seeking.
The 11-month gap in Section 4
Buried well past the headlines: from May 2025 through April 2026, Anthropic's blocking biological classifiers did not run on any traffic through its human feedback platforms. That covers roughly 133 million exchanges from about 50,000 contractors, vetted by outside vendors rather than by Anthropic.
An internal-only flag silently disabled both the blocking and the logging, so nothing was recorded for later review. A retroactive sweep using Claude Sonnet 5 flagged 1,197 high-risk transcripts. Manual review found no clearly concerning misuse.
Anthropic also retroactively downgraded its own February verdict from very low to low. A safety report that corrects the previous safety report is not a common document. Full detail here.
Young Americans do not trust the people running AI
A CNBC and Generation Labs survey of more than 1,000 Americans aged 18 to 34 found distrust running high across nearly every named executive:
| Executive |
Do not trust to act responsibly on AI |
| Alex Karp, Palantir |
81% |
| Peter Thiel |
79% |
| Mark Zuckerberg, Meta |
71% |
| Elon Musk, xAI |
70% |
| Sam Altman, OpenAI |
69% |
| Satya Nadella, Microsoft |
Only executive with net trust |
Beyond the personalities: 45 percent expect AI to hurt their careers, and 60 percent want the data-centre buildout slowed. That second number is the one with policy consequences, because data centres are the part of this industry that shows up in a local planning meeting.
DeepMind mapped the routes to superintelligence
Google DeepMind published a paper on August 14 framing the AGI-to-ASI transition around four technical pathways: continued scaling of compute, model size, data and test-time inference; algorithmic paradigm shifts beyond the transformer stack; recursive self-improvement where AI accelerates AI R&D; and multi-agent collective intelligence, where large populations of specialised agents coordinate into a superhuman group agent.
Read alongside Anthropic's saturated R&D benchmarks, pathway three is the one with a measurement problem attached. DeepMind separately published work on evaluating language models for harmful manipulation.
Deadlines still running
| Date |
What happens |
| Aug 31 |
Claude Sonnet 5 moves 2 to 3 dollars per million input, plus a tokenizer change adding 10 to 35 percent tokens on code |
| Aug 31 |
kimi-k2.5 and moonshot-v1 sunset, migrate to kimi-k3 |
| Early-mid Sept |
Grok 4.7 window, per Musk. No model ID at docs.x.ai yet |
| Oct 1 |
OpenAI vs Apple hearing |
| Oct 24 |
deepseek-chat and deepseek-reasoner deprecated |
What we are watching next
Every disclosure in the July and August incident wave arrived because a company chose to publish it. Anthropic's Long-Term Benefit Trust holds the power to compel external review of future risk reports. Whether it uses that power is the governance question this report raises and does not answer.
Second: whether any other frontier lab publishes a comparable self-assessment. Right now the company disclosing the most looks the worst, which is a bad incentive structure for everyone.
FAQ
Did an Anthropic model fail a safety test?
No. The rating moved because incident disclosures raised overall uncertainty. Anthropic states its underlying argument probably still supports the lower very low label, and says it hopes to return the rating there.
What is Anthropic Model 2?
An unreleased Mythos-class model, described as somewhat more capable than Mythos 5 and used heavily inside Anthropic. The report says there are no current plans to release it externally, and that Anthropic holds lower confidence in its capability estimates because the full evaluation suite has not been run.
Were customers affected by the classifier gap?
Per the report, no. The gap covered human-feedback contractor platforms rather than customer traffic, and the retroactive review found no clearly concerning misuse.
What does benchmark saturation mean here?
Anthropic's concrete task-based evaluations for automated AI R&D no longer register incremental capability gains. The measuring instrument has run out of range, which makes it harder to detect the acceleration it was built to catch.
When does Claude Sonnet 5 pricing change?
September 1, 2026. Input goes 2 to 3 dollars per million, output 10 to 15. A tokenizer change lands at the same time adding 10 to 35 percent more tokens on code, so the effective rise on coding workloads is larger than the sticker figure.