TUE, OCTOBER 06, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ General

A 501-Billion-Parameter Model You Cannot Download Yet

Reflection AI published Beam's full benchmarks - 80.9 on SWEBench Verified, 80.1 on Terminal Bench - under an Apache 2.0 commitment, with 23 billion active parameters out of 501 billion and a claim of comparable scores at three to four times less inference compute. The weights arrive later this month.

By AIToolsRecap October 6, 2026 6 min read 40 views
Home › Articles › General › Reflection's Beam: 501B Params, Apache 2.0, No ...

The Spec Sheet

PropertyBeam
Total parameters501 billion
Active parameters23 billion
ArchitectureSparse MoE, 52 layers
AttentionInterleaved local and global
Pretraining tokens23.8 trillion
Context (RL)256K
Context (effective)1M
LicenceApache 2.0
WeightsNot released - later this month

Benchmarks, all Reflection's own figures:

BenchmarkScore
SWEBench Verified80.9
Terminal Bench v2.180.1
SWE Bench Pro v2-Hard77.2
DeepSWE v1.144.4

23 Billion Active Is the Whole Product

A 501-billion-parameter model that only activates 23 billion per token costs roughly what a 23-billion model costs to serve, while drawing on the knowledge of a model twenty times larger. That ratio is the point of the architecture, and it is what produces the claim underneath the benchmarks.

Reflection says Beam reaches scores comparable to GLM-5.2 while using three to four times less inference compute. Read that as the headline rather than any individual benchmark number. Nobody serving a coding model at volume is short of options that score in the high seventies. They are short of options that score there cheaply.

The company positions it as competitive with GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic work, and concedes that frontier models like Kimi K3 lead on raw capability. That concession is worth noting - a vendor stating plainly which competitor beats it is doing something most launch posts avoid.

The Training Numbers Are Unusual

Two figures stand out from the disclosed methodology, and both are about filtering rather than scale.

Quality classifiers eliminated about 95% of raw internet tokens from the 23.8 trillion-token pretraining set, while recovering roughly 1.8 trillion high-quality tokens that conventional filters had missed. That is a claim about the pipeline being better at both throwing away and keeping - and the second half is the harder one. Code was tuned per language rather than treated as one category.

Then the reinforcement learning: 10,500 NVIDIA GB300 GPUs for four weeks, producing over 100 million rollouts across approximately 1.3 billion sandboxes and a million coding, agentic and STEM environments.

1.3 billion sandboxes is the number to sit with. That is the infrastructure cost of training an agentic model that nobody puts in the headline - not the GPUs, the execution environments. Every rollout needs somewhere isolated to run and something to check the result.

The Thing to Be Clear About

You cannot download Beam.

Reflection has published benchmarks, architecture, training methodology and an Apache 2.0 licence commitment. The weights, technical report, model card and developer artifacts arrive "later this month", after final red-teaming and evaluations. Access right now is a signup form.

That is a legitimate way to launch and it is also not the same thing as an open-weight release, however it gets reported. Until the weights land, every figure above is a first-party claim that nobody outside Reflection has been able to reproduce. Treat the announcement as an announcement.

Why Apache 2.0 Matters Here

If Beam ships under Apache 2.0 at this capability tier, it is the most permissive licence available on a frontier-class coding model.

Apache 2.0 means you can self-host it, modify it, put it inside a commercial product and ship that product, without a usage threshold, a revenue clause or a field-of-use restriction. Most open-weight releases at this scale carry at least one of those. For anyone who needs a capable coding model inside their own infrastructure - because of data residency, air-gapping, or simply because per-token pricing does not work at their volume - that combination is rare.

That is also why the delay matters more than it usually would. The licence is the differentiator, and the licence is not in effect until there is something to licence.

FAQ

Can I download Reflection Beam?

Not yet. Reflection says weights, the technical report, model card and developer artifacts arrive later in October 2026, after final red-teaming. Current access is through an early-access signup.

How big is Beam?

501 billion total parameters with 23 billion active per token - a sparse mixture-of-experts design across 52 layers, pretrained on 23.8 trillion tokens.

What does Beam score?

80.9 on SWEBench Verified, 80.1 on Terminal Bench v2.1, 77.2 on SWE Bench Pro v2-Hard and 44.4 on DeepSWE v1.1 - all Reflection's own figures, none independently verified.

What licence is it under?

Apache 2.0, which permits self-hosting, modification and commercial use without usage or revenue restrictions.

For where the closed models currently sit, see our October model rankings - and the reasons two major leaderboards disagree about them.

Tags
AI NewsCoding AIGenerative AI2026
⚑

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →
💡 AI Tools prompts
Prompt Guide
Best Claude AI Prompts for SEO (2026) — Content, Technical, and Comparison SEO
Claude Sonnet 5 and Opus 5 are strong for SEO work that requires writing quality, structured analysis, and long-form content generation. With 1M context, Claude can analyse an entire site's content structure, compare competing pages, and write complete article drafts in one session. These prompts cover the full SEO workflow: keyword research synthesis, content briefs, on-page optimisation, meta descriptions, technical audit interpretation, and comparison content that ranks above AI Overviews.
Get Prompts →
Prompt Guide
Best ChatGPT Prompts for SEO (2026) — GPT-5.6 and Browse
ChatGPT with GPT-5.6 Sol and Browse enabled is a capable SEO research tool — it can search the live web, analyse SERP results, and synthesise content briefs in a single session. GPT-5.6 Terra at $2.50/M offers a cost-efficient option for high-volume SEO content generation. These prompts are optimised for ChatGPT Plus with Browse, the ChatGPT Work product for larger projects, and the OpenAI API with web_search tool enabled.
Get Prompts →
Prompt Guide
Best Claude Opus 5 and Sonnet 5 Prompts for Writing (2026)
Claude Opus 5 and Sonnet 5 consistently produce the highest-quality long-form writing of any AI model in July 2026 — a lead documented across writing benchmarks and user testing since Claude 3 Opus. With 1M context and 128K output on Opus 5, Claude can write book chapters, complete reports, and long-form content without truncating. Sonnet 5 at $2/$10/M (intro through August 31) is the best value writing model available. These prompts are optimised for claude.ai Pro/Max, Claude Cowork, and the API.
Get Prompts →