The Benchmarks
| Benchmark | Gemini 4 Argon | Field |
| DeepSWE v1.1 | 77.9% | Claude Opus 5.5 74.2% · GPT-6 Astra 74.1% |
| CWE-bench v1 (vulnerability remediation) | 68% | tied for first |
| AutomationBench | 51.3% | ranked first |
| LVBench (video understanding) | 91.7% | state of the art |
A 3.7-point lead on DeepSWE over the two models it is competing against is a real margin on a coding benchmark, where the top models have been clustered within a point of each other for months.
One Million Output Tokens
The output limit goes from 64,000 to 1,000,000 tokens. That is the specification change that will actually alter what people build.
Input context has been large across the industry for a while; output has been the constraint. A model that can emit a million tokens in one response can produce an entire codebase, a full document set, or a long structured extraction without the chunking and stitching that currently eats engineering time and introduces errors at every seam.
Whether it holds quality across that length is the open question, and no one outside Google can answer it yet.
The Pricing Detail Everyone Will Skip
Launch pricing is $2 per million input and $10 per million output, with cached input discounted 95%.
Those are introductory rates. The standard rates are $4 and $20 — double.
This matters because of where it lands. Claude Sonnet 5.5 is $2 and $10. GPT-6.1 Sol is $2 and $10. Gemini 4 Argon is $2 and $10 for now.
So the convergence everyone is about to write about is partly an illusion. Two of those three prices are the actual price. The third is a trial.
If you are costing a migration, build the model on $4 and $20 and treat the introductory period as a discount rather than a rate. Google has not published how long it lasts.
Who Can Get It
Rolling out with priority to Google AI Ultra subscribers and paid API customers. Right now it is limited to trusted testers and cyber defenders through the Fairwind Program — the same gated route Google used for Gemini 3.8 Flash Cyber earlier this month.
That is now three labs in four weeks releasing frontier capability through a vetted-access programme first: Google with Fairwind, OpenAI with Daybreak Blue for Astra, and Anthropic routing higher-risk cyber work away from its newest Sonnet. Gated release has gone from unusual to standard practice without anyone announcing the shift.
Argon also follows the discontinuation of Gemini 3.5 Pro.
Where This Leaves the Three-Way Comparison
- On published coding benchmarks: Argon leads, by a margin that is small but larger than the usual noise between flagship releases.
- On price: level with Sonnet 5.5 and GPT-6.1 Sol today, double them later.
- On output length: not close. One million against anything currently shipping.
- On availability: last. Both competitors are generally available; Argon is gated.
The sensible move is to wait for the general release and benchmark it on your own workload. Every number above is Google's, published on launch day, with no independent verification yet.
FAQ
How much does Gemini 4 Argon cost?
$2 per million input tokens and $10 per million output during the introductory period, rising to $4 and $20 afterwards. Cached input is discounted 95%.
Is it better than Claude Opus 5.5?
On DeepSWE v1.1, Google reports 77.9% against 74.2% for Opus 5.5 and 74.1% for GPT-6 Astra. These are Google's own figures and have not been independently verified.
What is the output token limit?
One million, up from 64,000.
Who can use it?
Google AI Ultra subscribers and paid API customers get priority. Access is currently restricted to trusted testers and cyber defenders via the Fairwind Program.