GROK 4.6 — KEY FACTS (RELEASED AUGUST 12, 2026)
● Release date: August 12, 2026
● API price (under 200K tokens): $2/$6/M input/output — same as Grok 4.5
● API price (over 200K tokens): $4/$12/M — doubles, applies to entire request
● Cached input: $0.50/M
● Context window: 500,000 tokens (expanded from Grok 4.5)
● AA Intelligence Index: 61 — ties GPT-5.6 Sol Max, one below Fable 5 Max (62)
● CursorBench v3.2: 69.9% vs 66.7% for Grok 4.5
● APEX-Agents: Grok 4.6 leads GPT-5.6 Sol
● DeepSWE (repo-scale coding): GPT-5.6 Sol still leads
● Knowledge cutoff: February 1, 2026
● Input/output: Text + image input, text output only
● Available: xAI API · Grok Build · Cursor · OpenRouter · Vercel · Cloudflare
● Also launched Aug 11: Grok Bot (agent-teammate product)
● Next: Grok 4.7 (2.1T architecture, weeks away) · Grok 5 (before end 2026)
What Actually Changed From Grok 4.5
Per DEV Community's technical analysis, Grok 4.6 is not a new architecture — it is the same 1.5 trillion parameter base as Grok 4.5 with an upgraded post-training run. Three training ingredients drove the improvement: curated model-generated data targeting reasoning and advanced technical concepts; high-quality engineering datasets continuing the coding-heavy diet from real development sessions via Grok Build; and reinforcement learning in agentic environments covering knowledge work, general coding, web development, computer-aided design, and kernel optimisation. The behavioral result, per MarkTechPost's deep-dive, is that "on longer trajectories, xAI reports more self-testing and verification, with the model checking its own work before moving on." That is the specific capability gap over 4.5 — not raw intelligence score improvement, but sustained agent quality on long multi-step tasks.
As Kingy AI's launch review explains, the positioning is deliberate: "This is not pitched as a raw intelligence jump. It is a model built for long-running agents and ambitious interactive and visual work." The 500K context window supports this — Grok 4.5 had a shorter context, and the 4.6 expansion is directly targeted at the sustained-context use cases (codebase traversal, multi-document research, long agent trajectories) where context limits previously forced chunking.
The Benchmark Picture — Where It Leads and Where It Loses
| Benchmark | Grok 4.6 | GPT-5.6 Sol | Claude Fable 5 | Winner |
| AA Intelligence Index | 61 | 61 | 62 | Fable 5 (Grok ties Sol) |
| CursorBench v3.2 | 69.9% | — | — | Grok 4.6 (vs 4.5's 66.7%) |
| APEX-Agents | Leads | Behind | — | Grok 4.6 |
| DeepSWE (repo-scale coding) | Behind | Leads | — | GPT-5.6 Sol |
| BenchLM Agentic rank | #19 | — | — | Top 20 agentic |
Sources: APIdog benchmark breakdown, BenchLM August 13 data. xAI's table mixes public and self-reported competitor numbers — production teams should test on their own workloads.
The Pricing Trap — Read Before Using Long Context
Per Kingy AI's pricing analysis, "the headline API rate is $2 per million input tokens and $6 per million output tokens, but that shorthand is incomplete." The $2/$6/M rate applies only to requests with a prompt under 200,000 tokens. Once a prompt crosses that threshold, the rate doubles to $4/$12/M — and critically, this higher rate applies to all tokens in that request, not just the tokens above 200K. As Digital Applied's launch analysis notes, "Grok 4.6's real opportunity is therefore not 'stuff 500K tokens into every prompt.' It is using a stronger long-horizon model with disciplined context, caching, and compaction." If your prompt is 201K tokens, you pay $4/M on all 201K — not $2/M on the first 200K and $4/M on the marginal 1K.
The practical implication: Grok 4.6 is genuinely competitive at $2/$6/M for requests under 200K tokens. It is a different pricing story at $4/$12/M for long-context requests. At that rate, Claude Sonnet 5 ($2→$3/M from September 1) and even Claude Opus 5 ($5/M) become more competitive on a cost basis depending on output length. Per Digital Applied's analysis, the cached input rate ($0.50/M) is where the real savings lie for Grok 4.6 workloads — teams that aggressively cache stable system prompts and context will see the price advantage hold even on longer requests.
Who Should Use Grok 4.6
Grok 4.6 for: Long-running agent tasks (APEX-Agents leader), agentic coding in Cursor and Grok Build (CursorBench 69.9%), cost-sensitive API workloads under 200K tokens ($2/M), and teams that can aggressively cache context ($0.50/M cached). The frontier-tier performance at roughly one-fifth the output price of comparable models is the core value proposition.
Consider alternatives if: Repository-scale coding is your primary benchmark (GPT-5.6 Sol leads DeepSWE). Your prompts regularly exceed 200K tokens (long-context pricing doubles). You need Anthropic FLI C+ safety posture for enterprise procurement. You want the highest intelligence ceiling (Claude Fable 5 Max at 62 AA Index). Grok 4.5 hallucination concern (54% per AA) — verify independently on 4.6 before production deployment.
Cursor and Grok Build users: xAI is running a launch promo — 2× included usage in Cursor and Grok Build for the first week. No base usage figure published, so the multiplier is not directly quantifiable, but the model is live in both tools now.
What's Next — Grok 4.7 and Grok 5
Per Netalith's roadmap coverage, Elon Musk confirmed on X that Grok 4.7 is expected within weeks of 4.6's release, and a leak in early August pointed to 4.7 carrying a much larger 2.1 trillion parameter architecture — a genuine scale jump versus 4.6's post-training upgrade. Grok 5 is targeted before the end of 2026. If Grok 4.6 is the extraction run (how much can you improve from an existing base) and Grok 4.7 is the scale run (how much does a bigger base add), the August-September period at xAI will be the most significant model release period since Grok 4.0 launched in early 2026.
Sources: DEV Community technical analysis · Kingy AI pricing deep-dive · APIdog full benchmark breakdown · MarkTechPost training details · Digital Applied launch analysis · BenchLM data · Netalith roadmap · Related: Grok Build vs Claude Code → · AI model release tracker →