WED, SEPTEMBER 02, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ News

Nobody Publishes What a Limit Is in Tokens, So You Find the Ceiling by Hitting It

Google begins rolling out compute-specific flexible usage limits for Gemini Notebook consumer accounts today, Anthropic settles Claude Code weekly limits 17 percent below current levels on 14 September, and OpenAI restored a strict five-hour cap on Codex last week. Three vendors in one fortnight, all reaching the same conclusion about agent workloads, and none of them stating a limit in tokens.

By AIToolsRecap September 2, 2026 6 min read 18 views
Home Articles News Gemini Three Vendors Restructured Usage Limits in Two ...
THREE CHANGES, ONE FORTNIGHT

● Google, from today: compute-specific flexible usage limits for Gemini Notebook consumer accounts, web and mobile.

● Anthropic, 14 September: Claude Code weekly limits settle 17 percent below current levels.

● OpenAI, last week: restored a strict five-hour cap on Codex and Work tools.

● Common to all three: none publishes what a limit is in tokens.

What Google is doing

Google announced on 28 August that it will roll out compute-specific, flexible usage limits for Gemini Notebook consumer accounts on web and mobile, starting 2 September.

The word doing the work is compute-specific. A limit tied to compute consumed rather than requests made is a different model from counting messages — it means a long, tool-heavy session costs more of your allowance than a short question, which is closer to what the work actually costs to serve.

The pattern

Vendor Change Direction
Google Compute-specific flexible limits, from 2 September Restructure
Anthropic Weekly limits permanent at 25% above old baseline, from 14 September Down 17% on today
OpenAI Five-hour cap restored on Codex and Work tools Tightened
WHAT THIS IS ACTUALLY ABOUT

Flat-rate subscriptions were priced to acquire users during a period when nobody knew what heavy use looked like. Agents changed that — a single unattended session can consume more than a month of chat.

Three vendors adjusting within a fortnight is not coordination. It is three companies reaching the same conclusion about the same workload at roughly the same time.

The thing none of them do

None of the three publishes what a limit is in tokens.

Anthropic gives multipliers — Max 5x is five times Pro. OpenAI publishes five-hour allowances with a footnote that additional weekly limits may apply, and no figures for those. Google is introducing compute-specific limits without saying what a unit of compute is.

The practical consequence is the same everywhere: you discover your ceiling by hitting it. That makes capacity planning guesswork, and it is the single most common complaint across all three products.

The vendors have a real reason — consumption genuinely varies with model, context length and tool use, so a token figure would mislead as often as it helped. That explanation is honest and it does not make planning any easier.

What to do about it

If you are... Do this
On Claude Pro or Max Measure now, before 14 September. Above 83 percent on the weekly bar means the new ceiling bites
Using Gemini Notebook Watch what compute-specific means in practice this week. Long sessions may cost more than before
On Codex Free or Go You stop when capped — neither tier can buy credits
Budgeting for Q4 Assume limits tighten rather than loosen. Every change this fortnight went one way
Running agents unattended This is the workload the changes are aimed at. Cap context and diff size before paying for more

Sources

FAQ

What is changing with Gemini Notebook?

Google is rolling out compute-specific, flexible usage limits for consumer accounts on web and mobile from 2 September 2026, replacing a simpler allowance model.

What does compute-specific mean?

Consumption is measured by compute used rather than requests made, so a long tool-heavy session draws more of your allowance than a short question.

Are all three vendors reducing limits?

Anthropic and OpenAI have both tightened. Google is restructuring rather than explicitly reducing, and what it means in practice will be visible this week.

Why does none of them publish token limits?

Because consumption varies with model, context length and tool use, so a single figure would mislead. The consequence is that you find your ceiling by reaching it.

What should I check first?

Your own usage history on whichever product you rely on. Every vendor exposes a usage view, and two weeks of data before a change is worth more than any published figure.

Tags
GoogleGeminiAnthropicClaude CodeOpenAICodexUsage LimitsPricing2026

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →