THREE CHANGES, ONE FORTNIGHT
● Google, from today: compute-specific flexible usage limits for Gemini Notebook consumer accounts, web and mobile.
● Anthropic, 14 September: Claude Code weekly limits settle 17 percent below current levels.
● OpenAI, last week: restored a strict five-hour cap on Codex and Work tools.
● Common to all three: none publishes what a limit is in tokens.
What Google is doing
Google announced on 28 August that it will roll out compute-specific, flexible usage limits for Gemini Notebook consumer accounts on web and mobile, starting 2 September.
The word doing the work is compute-specific. A limit tied to compute consumed rather than requests made is a different model from counting messages — it means a long, tool-heavy session costs more of your allowance than a short question, which is closer to what the work actually costs to serve.
The pattern
| Vendor |
Change |
Direction |
| Google |
Compute-specific flexible limits, from 2 September |
Restructure |
| Anthropic |
Weekly limits permanent at 25% above old baseline, from 14 September |
Down 17% on today |
| OpenAI |
Five-hour cap restored on Codex and Work tools |
Tightened |
WHAT THIS IS ACTUALLY ABOUT
Flat-rate subscriptions were priced to acquire users during a period when nobody knew what heavy use looked like. Agents changed that — a single unattended session can consume more than a month of chat.
Three vendors adjusting within a fortnight is not coordination. It is three companies reaching the same conclusion about the same workload at roughly the same time.
The thing none of them do
None of the three publishes what a limit is in tokens.
Anthropic gives multipliers — Max 5x is five times Pro. OpenAI publishes five-hour allowances with a footnote that additional weekly limits may apply, and no figures for those. Google is introducing compute-specific limits without saying what a unit of compute is.
The practical consequence is the same everywhere: you discover your ceiling by hitting it. That makes capacity planning guesswork, and it is the single most common complaint across all three products.
The vendors have a real reason — consumption genuinely varies with model, context length and tool use, so a token figure would mislead as often as it helped. That explanation is honest and it does not make planning any easier.
What to do about it
| If you are... |
Do this |
| On Claude Pro or Max |
Measure now, before 14 September. Above 83 percent on the weekly bar means the new ceiling bites |
| Using Gemini Notebook |
Watch what compute-specific means in practice this week. Long sessions may cost more than before |
| On Codex Free or Go |
You stop when capped — neither tier can buy credits |
| Budgeting for Q4 |
Assume limits tighten rather than loosen. Every change this fortnight went one way |
| Running agents unattended |
This is the workload the changes are aimed at. Cap context and diff size before paying for more |
Sources
FAQ
What is changing with Gemini Notebook?
Google is rolling out compute-specific, flexible usage limits for consumer accounts on web and mobile from 2 September 2026, replacing a simpler allowance model.
What does compute-specific mean?
Consumption is measured by compute used rather than requests made, so a long tool-heavy session draws more of your allowance than a short question.
Are all three vendors reducing limits?
Anthropic and OpenAI have both tightened. Google is restructuring rather than explicitly reducing, and what it means in practice will be visible this week.
Why does none of them publish token limits?
Because consumption varies with model, context length and tool use, so a single figure would mislead. The consequence is that you find your ceiling by reaching it.
What should I check first?
Your own usage history on whichever product you rely on. Every vendor exposes a usage view, and two weeks of data before a change is worth more than any published figure.