The Spec Sheet
| Property | Beam |
| Total parameters | 501 billion |
| Active parameters | 23 billion |
| Architecture | Sparse MoE, 52 layers |
| Attention | Interleaved local and global |
| Pretraining tokens | 23.8 trillion |
| Context (RL) | 256K |
| Context (effective) | 1M |
| Licence | Apache 2.0 |
| Weights | Not released - later this month |
Benchmarks, all Reflection's own figures:
| Benchmark | Score |
| SWEBench Verified | 80.9 |
| Terminal Bench v2.1 | 80.1 |
| SWE Bench Pro v2-Hard | 77.2 |
| DeepSWE v1.1 | 44.4 |
23 Billion Active Is the Whole Product
A 501-billion-parameter model that only activates 23 billion per token costs roughly what a 23-billion model costs to serve, while drawing on the knowledge of a model twenty times larger. That ratio is the point of the architecture, and it is what produces the claim underneath the benchmarks.
Reflection says Beam reaches scores comparable to GLM-5.2 while using three to four times less inference compute. Read that as the headline rather than any individual benchmark number. Nobody serving a coding model at volume is short of options that score in the high seventies. They are short of options that score there cheaply.
The company positions it as competitive with GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic work, and concedes that frontier models like Kimi K3 lead on raw capability. That concession is worth noting - a vendor stating plainly which competitor beats it is doing something most launch posts avoid.
The Training Numbers Are Unusual
Two figures stand out from the disclosed methodology, and both are about filtering rather than scale.
Quality classifiers eliminated about 95% of raw internet tokens from the 23.8 trillion-token pretraining set, while recovering roughly 1.8 trillion high-quality tokens that conventional filters had missed. That is a claim about the pipeline being better at both throwing away and keeping - and the second half is the harder one. Code was tuned per language rather than treated as one category.
Then the reinforcement learning: 10,500 NVIDIA GB300 GPUs for four weeks, producing over 100 million rollouts across approximately 1.3 billion sandboxes and a million coding, agentic and STEM environments.
1.3 billion sandboxes is the number to sit with. That is the infrastructure cost of training an agentic model that nobody puts in the headline - not the GPUs, the execution environments. Every rollout needs somewhere isolated to run and something to check the result.
The Thing to Be Clear About
You cannot download Beam.
Reflection has published benchmarks, architecture, training methodology and an Apache 2.0 licence commitment. The weights, technical report, model card and developer artifacts arrive "later this month", after final red-teaming and evaluations. Access right now is a signup form.
That is a legitimate way to launch and it is also not the same thing as an open-weight release, however it gets reported. Until the weights land, every figure above is a first-party claim that nobody outside Reflection has been able to reproduce. Treat the announcement as an announcement.
Why Apache 2.0 Matters Here
If Beam ships under Apache 2.0 at this capability tier, it is the most permissive licence available on a frontier-class coding model.
Apache 2.0 means you can self-host it, modify it, put it inside a commercial product and ship that product, without a usage threshold, a revenue clause or a field-of-use restriction. Most open-weight releases at this scale carry at least one of those. For anyone who needs a capable coding model inside their own infrastructure - because of data residency, air-gapping, or simply because per-token pricing does not work at their volume - that combination is rare.
That is also why the delay matters more than it usually would. The licence is the differentiator, and the licence is not in effect until there is something to licence.
FAQ
Can I download Reflection Beam?
Not yet. Reflection says weights, the technical report, model card and developer artifacts arrive later in October 2026, after final red-teaming. Current access is through an early-access signup.
How big is Beam?
501 billion total parameters with 23 billion active per token - a sparse mixture-of-experts design across 52 layers, pretrained on 23.8 trillion tokens.
What does Beam score?
80.9 on SWEBench Verified, 80.1 on Terminal Bench v2.1, 77.2 on SWE Bench Pro v2-Hard and 44.4 on DeepSWE v1.1 - all Reflection's own figures, none independently verified.
What licence is it under?
Apache 2.0, which permits self-hosting, modification and commercial use without usage or revenue restrictions.
For where the closed models currently sit, see our October model rankings - and the reasons two major leaderboards disagree about them.