WHAT MUSK SAID, AND WHERE
● Grok 4.8: 2.5 trillion parameters, new C++ training stack, finishes training this week then starts RL.
● Grok 4.9: likely Astra or Fable class, per Musk.
● Grok 5: asked how close 4.8 gets to AGI, he replied "That will be Grok 5".
● The source for all of it: posts on X. xAI has published nothing.
The model before this one has not shipped
Grok 4.7 was scheduled for 12 September. It is 14 September and there is no model ID at docs.x.ai, no price and no benchmark card.
On 11 September Musk said it needed a few more days to cook. The reason he gave is the most technically interesting thing in this whole sequence: reinforcement learning may have penalised response length too heavily, leaving the model prone to giving up on hard tasks early and not checking its work rigorously.
THAT ADMISSION IS WORTH MORE THAN THE ROADMAP
A model trained to be concise learning to quit early is a real and specific failure mode. It is the kind of detail labs usually keep internal, and it tells you something about how these systems are tuned that no benchmark does.
It is also the reason to be careful with the rest. If 4.7 slipped on a problem found in RL, 4.8 has not started RL yet.
The roadmap, and what is behind each number
| Model |
Claim |
Source |
| Grok 4.7 | ~2.1T, SpaceX engineering data, on par with Opus 5.0 | Musk on X. Unreleased |
| Grok 4.8 | 2.5T, C++ stack, "noticeable improvement" | Musk on X, 14 Sept |
| Grok 4.9 | Astra or Fable class | Musk on X |
| Grok 5 | AGI | Musk on X |
| A 3T run | Upgraded training software, cleaner data | Musk on X, no name yet |
Four unreleased models and a fifth training run, announced in about thirty-six hours. Not one of them appears in an xAI publication.
The C++ stack is the substantive part
xAI has been rewriting its training and inference stack in C/C++ since June, aimed at getting more out of NVIDIA GB300 hardware in Colossus. Grok 4.8 is the first model Musk has described as trained on it.
He also contrasted it with an earlier 2.1T run built with Jax that had mistakes only corrected mid-run — which, if accurate, means the previous model was trained through errors that were patched while it was learning. That is a more meaningful claim than any parameter count, because it goes to whether the training was clean rather than how big it was.
What parameter counts do and do not tell you
- 2.5 trillion is large, and size is not capability. K2 Horizon released six models with full training data and none of them lead on benchmarks by parameter count alone.
- Recent research suggests data matters more. Analysis of 2019 to 2025 found roughly three times more compute-efficiency gain came from data improvements than from model architecture.
- RL is where this line has stumbled. 4.7 is late because of a reinforcement learning problem, and 4.8 has not started RL yet.
- Nothing here is testable. No weights, no API, no benchmark card, no price.
On the AGI claim
Asked how close Grok 4.8 would come to AGI, Musk replied: "That will be Grok 5."
There is no agreed definition of AGI, no benchmark that resolves it, and no way for anyone outside xAI to test the claim when the model arrives. Treat it as a statement of ambition rather than a prediction with a date attached — which is how it was made.
What to do with this
| If you are... |
The read |
| Deciding whether to wait for 4.7 | Do not. It has missed late August, early September and 12 September |
| On SuperGrok deciding about an upgrade | Model upgrades have reached all paid tiers, so the version is not the variable |
| Building on the Grok API | Nothing changes until a model ID appears at docs.x.ai |
| Comparing labs | Astra and Fable 5.1 shipped with published prices and cards. That is the difference |
Every Grok version mapped to a model you can already buy →
Sources
FAQ
What is Grok 4.8?
Per Musk on X, a 2.5 trillion parameter model trained with xAI's new C++ software stack, finishing training this week before entering reinforcement learning. xAI has published nothing about it.
Has Grok 4.7 been released?
No. It was scheduled for 12 September and remains unreleased, with no model ID, price or benchmark card published.
Why was Grok 4.7 delayed?
Musk said reinforcement learning may have penalised response length too heavily, leaving the model prone to giving up on hard tasks early and not checking its work rigorously.
Will Grok 5 be AGI?
Musk said "That will be Grok 5" when asked how close 4.8 would come. There is no agreed definition of AGI and no benchmark that resolves the question, so treat it as ambition rather than a dated prediction.
Is 2.5 trillion parameters a lot?
It is large, but size is not capability. Recent analysis suggests data quality contributed roughly three times more compute-efficiency gain than architecture between 2019 and 2025.
When can I use any of these?
Unknown. No release date has been given for 4.7, 4.8, 4.9 or 5, and none appears in xAI documentation.