THE 60-SECOND VERSION
● Anthropic gave METR wide access to millions of evaluation and production transcripts.
● OpenAI shipped GPT-Live-1, full-duplex voice through the API.
● 2 days to Claude Code weekly limits settling 17 percent lower.
Anthropic opens its transcripts
After disclosing four real-system breaches during evaluations, Anthropic agreed to give METR wide access for an independent investigation — scanning millions of evaluation and production transcripts.
Transcripts are the most guarded asset any lab holds. An outside body reading them can find patterns the internal team missed, including ones the lab would rather not publish. That is what makes the commitment mean something.
The questions that decide how much it is worth have not been answered publicly: does METR publish independently, can Anthropic review before release, and what was excluded from the access.
Five things to watch, including the one enterprise buyers will ask →
OpenAI ships full-duplex voice
GPT-Live-1 is available through the API. It listens and speaks simultaneously rather than taking turns, handles interruptions natively, and delegates deeper reasoning to backend models and tools.
The point is removing the handoffs. Most voice systems chain recognition, then a language model, then synthesis — each boundary adding latency and each one a place things break. Interruption is the clearest case: ordinary in human conversation, an exception in turn-based software.
What you would build with it, and what to check first →
Also this week
- Magic says its pretraining recipe is more than ten times more compute-efficient than leading open-weight base models — the argument being that small labs can only compete on algorithmic efficiency.
- New research on pretraining finds that between 2019 and 2025, roughly three times more compute-efficiency gain came from data improvements than from model architecture, with small models benefiting most from data quality.
- CNBC reports Chinese labs accelerating — Moonshot, DeepSeek and Alibaba converging on frontier developers for model quality, coding ability and cost efficiency.
- Google Threat Intelligence flagged autonomous AI credential harvesting as a rising category.
The rest of the month
| Date | What happens |
| Sept 14 | Claude Code weekly limits settle 17 percent below current levels |
| Sept 29 | OpenAI DevDay, San Francisco |
| Oct 1 | OpenAI vs Apple hearing |
| Oct 24 | deepseek-chat and deepseek-reasoner deprecated |
| Nov 12 | OpenAI models leave Cursor |
Still open
FAQ
What did Anthropic agree with METR?
Wide access for an independent investigation, including scanning millions of evaluation and production transcripts, following disclosure of four real-system breaches during evaluations.
What is GPT-Live-1?
An OpenAI voice model in the API supporting full-duplex operation — listening and speaking simultaneously, handling interruptions natively, and delegating deeper reasoning to backend models.
What is the next confirmed AI deadline?
14 September, when Claude Code weekly limits settle at 25 percent above the pre-May baseline — 17 percent below current levels.