Reflection releases 501B-parameter Beam model
Beam is a sparse mixture-of-experts model with 23B active parameters, a reported 1M-token effective context, and Apache 2.0 weights. Independent benchmarking is beginning.

TL;DR
- Beam is Reflection’s first 501B sparse MoE with 23B active parameters, a specification reported in kimmonismus's Beam launch post, for coding, reasoning, and agentic work.
- Beam’s published scorecard reaches 80.9 on SWE-bench Verified and 80.1 on Terminal-Bench 2.1, with kimmonismus's launch post showing stronger and weaker results across the comparison set.
- Reflection’s 3–4x inference-efficiency claim uses estimated forward-pass compute, a method rohanpaul_ai's breakdown describes as excluding prefill, attention, and serving overhead.
- Reflection’s commercial pitch is control and customization: kimmonismus's early report described systems built on a company’s own data and compute, while WesRoth's summary said the Apache 2.0 weights were still forthcoming.
- Independent benchmarking is beginning, with ArtificialAnlys's update saying Artificial Analysis has access but has not yet published a scorecard.
Reflection’s official announcement includes an RL system that pushed new weights to its inference fleet in a median 12 seconds and handled 71 inference incidents without ending the training job. The model’s context accounting carries two reported numbers, while rohanpaul_ai's breakdown cites a 1M-token effective context and Reflection’s technical description gives a 256K maximum for the RL run. ArtificialAnlys's update says outside benchmarking has started before the public weights arrive.
Model shape
Beam is a sparse mixture-of-experts model, with Reflection positioning it for coding, reasoning, and tool-using agent workloads in its official announcement. The supplied launch graphic provides the compact architecture summary:
- Total parameters: 501 billion, according to kimmonismus's launch post and the official announcement.
- Active parameters: 23 billion per token.
- Modality: Text-only, although the official demos show it working with other modalities represented as text.
- Context: rohanpaul_ai's breakdown reports a 1M-token effective context, while Reflection says the high-compute RL run used a 256K maximum context.
Reflection describes Beam’s edge as inference efficiency rather than the highest raw capability. Its own announcement says Kimi K3 remains ahead on raw capability while Beam targets a lower-compute operating point.
Benchmark scorecard
Reflection’s published scorecard is a mixed comparison, not a single overall ranking. RuntimeWire likewise describes the results as company-published comparisons across selected coding and reasoning tests in its launch analysis.
- SWE-bench Verified: 80.9, above Inkling at 77.6 and Nemotron 3 Ultra at 70.7 in the original graphic, while other model results were not reported there.
- SWE-bench Pro v1: 65.5, below Qwen 3.8 Max at 67.7 and above GLM 5.2 at 62.1.
- Terminal-Bench 2.1: 80.1, just below GLM 5.2 at 81.0 and behind Qwen 3.8 Max at 86.6.
- DeepSWE v1.1: 44.4, below Qwen 3.8 Max at 51.0.
- HLE no tools: 36.2, below GLM 5.2 at 40.5 and Qwen 3.8 Max at 43.6.
CritPT AA is inconsistent across the launch artifacts: kimmonismus's launch post shows 15.6, while the current official table lists 16.3. The gap is small, but benchmark tables are only useful when the published values stay stable.
Compute claim
Reflection says Beam reaches reasoning scores comparable to GLM 5.2 with 3–4x less inference compute. The official methodology defines that comparison as estimated generation forward FLOPs per attempt:
- Formula: approximately 2 × active parameters × mean generated tokens per attempt.
- Included: reasoning tokens and final-answer tokens.
- MoE treatment: active parameters, not total parameters.
- Excluded: prompt prefill, context-dependent attention, and serving overhead.
Those exclusions are explicit in rohanpaul_ai's breakdown, which frames the number as an estimate of model compute rather than a measured customer serving cost. RuntimeWire makes the same distinction in its analysis.
Training run
Reflection says Beam was pretrained on 23.8 trillion tokens, then subjected to a four-week high-compute RL run that generated more than 100 million rollouts on 10,500 NVIDIA GB300 GPUs. The official report describes the run as a central scaling axis rather than a short post-training pass.
- Pretraining: under four weeks on 6,144 NVIDIA GB300 NVL72 GPUs.
- RL: four weeks on 10.5K GB300 GPUs, with a 256K maximum context and approximately 1.3 billion sandboxes used for training and grading.
- Environment pool: nearly one million coding, agentic, STEM, web-search, and tool-use environments.
- Optimization: asynchronous policy gradients, including training from rollouts more than a day old.
The infrastructure report says pretraining reached 92.3% goodput toward the end after nine semi-automatic rewinds. eliebakouch's comment estimated roughly 12% BF16 MFU using an 80% goodput assumption and questioned the utilization, while eliebakouch's reply said Reflection had intentionally underestimated goodput to produce a higher-bound MFU estimate.
RL infrastructure
The useful technical detail sits below the headline GPU counts. Reflection’s infrastructure section reports:
- Rollout concurrency: 110K concurrent rollouts on average.
- Weight distribution: new weights reached the inference fleet in about 12 seconds at the median; hierarchical distribution cut cross-rack traffic by 75% and made fleet-wide adoption 2.2x faster.
- Failure recovery: 71 inference incidents were handled without terminating training; capacity recovered in a median eight minutes, with lost capacity equal to 0.02% of serving GPU-minutes.
- Sandbox scale: up to 170K concurrent sandboxes, more than one billion sandbox creation requests, and 90% of new sandboxes ready in under 10 seconds.
- Batch packing: dynamic packing kept training batches 99.99% full on average while per-GPU throughput stayed within 1.5% as mean rollout length grew almost 70%.
A follow-up from teortaxesTex's follow-up withdrew a simple three-to-five-minute estimate for each sandbox, noting that Reflection’s count may include short grading jobs and that other systems use different sandbox-counting conventions.
Open release
Beam was announced before its weights were downloadable. Reflection’s official announcement says final red-teaming and evaluations are still underway, with early access available by sign-up and the weights, technical report, model card, and developer artifacts planned for later in October.
The planned distribution has three concrete pieces:
- License: Apache 2.0 weights, as WesRoth's summary also described when the model was announced.
- Builds: FP8 and NVFP4 variants are expected alongside the weights, according to rohanpaul_ai's breakdown.
- Ecosystem: Ollama welcomed the model in ollama's post and said more U.S. models were coming to its runtime; danielhanchen's reply said Unsloth would prepare quantizations.
Outside testing
Artificial Analysis says it has access and is independently benchmarking Beam, ArtificialAnlys's update but the post does not publish a score. At launch, Turing Post’s analysis said the model was not public enough for hands-on testing and added six Qwen and DeepSeek results from Mercor to Reflection’s comparison, including 83.6 for Qwen 3.8 Max and 82.9 for DeepSeek V4.1 Flash on SWE-bench Verified. The analysis also warns that evaluation settings differ.