THE SHORT VERSION
● Size: 1.05 trillion total parameters, 49B active per token (~4.7% activation)
● Context: 1 million tokens, native image input via a 1.6B-parameter vision encoder
● Price: $1.36 / $4.18 per million input / output tokens; $0.14 cached input
● Available: API public preview since 6 October 2026
● Weights: promised by end of October. Not released. No licence named yet.
Mistral shipped Large 4 on 6 October 2026, under the internal nickname Le Chonk. It is the largest model the company has released, and the pricing is the most interesting thing about it.
What is actually in it
Large 4 is a granular mixture-of-experts model: 1.05 trillion parameters in total, of which roughly 49 billion activate for any given token. That is about 4.7% of the weights doing work at any moment, which is how a trillion-parameter model gets served at these prices.
It is a hybrid instruct-and-reasoning model rather than two separate checkpoints, so there is no reasoning variant to switch to. Vision comes from a 1.6-billion-parameter encoder with native image input. Context is 1 million tokens.
Mistral trained it on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacentres, on data spanning more than 160 languages including every official EU language. For teams with data-residency requirements inside the EU, that training location is the detail that matters more than any benchmark.
Pricing, against what you are probably using
| Model |
Input / 1M |
Output / 1M |
Context |
| Mistral Large 4 |
$1.36 |
$4.18 |
1M |
| Gemini 4 Argon (intro) |
$2.00 |
$10.00 |
1M out |
| Gemini 4 Argon (list) |
$4.00 |
$20.00 |
1M out |
| GPT-6.1 Sol |
$2.00 |
$10.00 |
1M |
Output tokens are where API bills actually come from, and Large 4 is under half the price of both. Cached input at $0.14 per million makes long-document and repeated-context workloads cheaper again.
Benchmarks
Mistral published these at launch:
- Cybench 93%
- CyberGym-E2E 82%
- DeepSWE v1.1 61.7%
- SWE-Atlas-QnA 59.4%
- Terminal-Bench 4.0 28.3%
- Lakera B3 93.3% resistance
- KORABench 1.691 / 2.0
- Surge AI human evaluation 3.74 / 5.0
The security numbers are the standouts. Terminal-Bench at 28.3% is the one to note before you build an agent on this: long-horizon terminal work is still where every model struggles, and Large 4 is not the exception.
The open-weight claim needs an asterisk
WHAT YOU CAN AND CANNOT DO TODAY
● Can: call it through the API, in public preview, at the prices above
● Cannot: download the weights. They are promised by end of October 2026
● Unknown: the licence. Mistral has not named one
Large 4 is being discussed as an open-weight model. At the time of writing it is not one. The weights are promised by the end of October and the licence has not been announced, which means nobody can yet say whether it will be usable commercially, for fine-tuning, or for redistribution.
This is the second model in a week to launch on that pattern — announced as open, shipped as an API. If your plan depends on self-hosting, the model does not exist for you yet. Judge it when the weights land and the licence file is readable.
Who should actually try it
Use it if you are paying GPT-6.1 Sol or Gemini 4 Argon rates for high-output-volume work, or you need EU-trained, EU-hosted inference for compliance reasons, or you work in languages outside the usual ten.
Wait if you need the weights, since they are not out; or you are building long-horizon terminal agents, where that 28.3% is the number that will bite you.
For a broader look at what open-weight models can and cannot do on ordinary hardware, see our writeup on Gemma 4 and open models on everyday devices.
FAQ
Is Mistral Large 4 open source?
Not yet, and open-weight is the more accurate term in any case. Mistral has promised the weights by the end of October 2026 and has not announced a licence. Until both arrive, it is an API-only model.
How much does Mistral Large 4 cost?
$1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million.
What does 49B active parameters mean?
It is a mixture-of-experts model. Of the 1.05 trillion parameters stored, only about 49 billion are used to produce each token. That is why a trillion-parameter model can be served at under $5 per million output tokens.
Can it read images?
Yes. A 1.6-billion-parameter vision encoder handles native image input, inside the same 1-million-token context window.
Where was it trained?
On 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacentres, across 160+ languages including all official EU languages.
Sources