Reflection announces Beam, a 501B open-weight MoE model

On October 5, 2026, Reflection announced Beam, its first open-weight model. Beam is a sparse Mixture-of-Experts model with 501 billion total parameters and 23 billion active, built for coding, reasoning and agentic workloads. The weights are not out yet. Beam is in final red-teaming and evaluations, early-access sign-up is open, and Reflection says it will release the weights, a technical report, a model card and developer artifacts later this month.
Reflection puts Beam's capabilities down to two investments. The first is pretraining on 23.8 trillion tokens from the web and proprietary licensed datasets; Reflection says the resulting base model matches or outperforms available similar-sized open base models. The second is a high-compute reinforcement learning (RL) run that generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over four weeks, with a maximum context length of 256K tokens. Training and grading used about 1.3 billion sandboxes, and Reflection says it sourced one million coding, agentic and STEM environments to sustain the run. It believes this is one of the largest RL runs by any open lab to date, and says capabilities kept improving as RL compute grew, with no sign of a plateau.
On capability, Reflection claims Beam advances the Western open-weight frontier. It calls the model competitive with larger open models such as GLM-5.2 and approaching Qwen 3.8-Max on coding and agentic tasks. Frontier open models like Kimi K3 remain ahead on raw capability; Beam's stated advantage is efficiency at inference time. On advanced reasoning benchmarks it reports scores comparable to GLM-5.2 while using 3 to 4 times less inference compute, and says the gap widens against 2T+ parameter models like Qwen 3.8-Max. The benchmark figures themselves sit in a chart, not in the article text.
The RL method is asynchronous policy gradients. Reflection says policy staleness becomes a major source of instability at scale, because long rollouts contain tokens from several checkpoints, and that numerical mismatch between training and inference engines makes it worse. It developed new algorithms to keep learning stable, including when learning from interactions generated more than a day earlier. A controllable length penalty rewards correct solutions while discouraging wasted tokens. Early in RL, performance rose as completions got shorter; later, as agentic skill grew, completions lengthened again with further gains. Users can set the trade-off through a reasoning effort parameter: lower settings give shorter responses, higher ones allow longer reasoning.
Reflection also reports transfer between domains. During a phase trained on reasoning, software engineering and terminal tasks, it saw consistent gains in browsing even though no browsing tasks were in the RL mixture. Given web access, Beam learned on its own to search for and query other LLMs and to use OCR APIs to read documents. Demos include a live NYC subway dashboard built from public data, interactive applications, and a fine-tuning notebook for the latest and smallest Gemma-4 model on a Text2SQL task. Beam is text-only, though it can use other modalities when they are represented as text.
The environment pool held nearly one million tasks, built mainly through synthetic data pipelines plus proprietary vendor and open-source data. Reflection filtered them for difficulty (neither always solvable nor impossible) and quality (not underspecified, misleading, guessable or hackable), then tested them through RL to feed the next round. It says compromises in data quality caused capability plateaus and other training problems.
The infrastructure section lists specifics. Average concurrent rollouts were 110K. Inference-to-training GPU ratios ran from 3.9:1 to 5.4:1, and the trainer was resized across five GPU mesh configurations without losing state. New weights reached the inference fleet in a median of about 12 seconds; hierarchical distribution (across racks over RoCE, then locally over NVLink) cut cross-rack traffic by 75% and made fleet-wide adoption 2.2 times faster than every replica pulling directly. 71 inference incidents were handled without stopping the training job, with capacity recovering in a median of eight minutes and lost capacity at 0.02% of elapsed serving GPU-minutes. The sandbox platform supported up to 170K concurrent sandboxes, processed more than one billion creation requests across over 20 clusters, two clouds and four regions, and had 90% of new sandboxes ready in under 10 seconds. Dynamic packing kept training batches 99.99% full on average, holding per-GPU trainer throughput within 1.5% as mean rollout length grew almost 70%. Per-token records allowed consistency checks between training and inference, and independent judges re-screened passing solutions for verifier exploits.
On pretraining, Reflection trained a series of progressively bigger models to verify its scaling recipe, using in-house code and web validation sets and decontaminating training data against them; the final Beam Base matched its predicted performance. The architecture combines interleaved local and global attention, fine-grained routed experts, a controlled residual stream and multiple forms of load balancing. For expert balance it builds on auxiliary-loss-free load balancing and adds cosine decay of expert-bias updates to reduce routing perturbations late in training, plus sequence-level balancing; it says the final base has almost-perfect uniform expert utilization. For the residual stream it describes a depth-based scaling approach combined with SandwichNorm, elementwise attention gating and FP32 residual accumulation. The source text breaks off partway through this section.
Key facts
- Beam is Reflection's first open-weight model: a sparse Mixture-of-Experts with 501 billion total and 23 billion active parameters, aimed at coding, reasoning and agentic work.
- Pretraining used 23.8 trillion tokens; the RL run produced over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over four weeks.
- Reflection says Beam is competitive with GLM-5.2 and approaching Qwen 3.8-Max on coding and agentic tasks, with 3 to 4 times less inference compute than GLM-5.2 on advanced reasoning benchmarks; Kimi K3 stays ahead on raw capability.
- Weights, technical report, model card and developer artifacts are promised later this month; early-access sign-up is open and final red-teaming is under way.
- A reasoning effort parameter lets users trade shorter responses for longer reasoning; Beam is text-only.
Why it matters
Beam is Reflection's first open-weight release, and the company frames it as advancing the Western open-weight frontier. The pitch is not top raw capability: Reflection concedes that frontier open models like Kimi K3 stay ahead there. The claim is efficiency, with comparable scores to GLM-5.2 on advanced reasoning benchmarks at 3 to 4 times less inference compute. The announcement also details an unusually large RL run, which Reflection believes is one of the largest by any open lab, and says gains had not plateaued when it ended.
Who it affects
Reflection aims Beam at enterprise coding and agentic workloads, calling it a workhorse model that delivers strong capability at lower cost. Teams that build coding agents or run open-weight models themselves are the obvious audience. Researchers working on large-scale asynchronous RL may also find the infrastructure details useful, from weight distribution to sandbox scaling.
How to use it
Not yet downloadable. Reflection has opened an early-access sign-up, and says the weights, technical report, model card and developer artifacts will follow later this month. No exact date is given. Once available, the reasoning effort parameter lets you choose shorter responses at lower settings or longer reasoning at higher ones to match the task and compute budget. Beam takes text only, though other modalities can be fed in if represented as text. No licence for the weights is stated in the visible text.
How solid is it
This is Reflection's own announcement, and every comparison in it is the company's own. No independent evaluation or third-party verification is given. The benchmark scores are in a figure, so the article text itself carries no per-benchmark numbers. The infrastructure and training figures (rollouts, GPUs, sandboxes, incident counts) are specific and stated plainly, but they are self-reported. Claims such as one of the largest open-lab RL runs are framed by Reflection as belief.
Risks and caveats
The weights are not out, so none of the claims can yet be tested by outsiders. Beam is still in final red-teaming and evaluations, and the release is stated as intent for later this month. The model trails Kimi K3 on raw capability by Reflection's own account, and approaches rather than matches Qwen 3.8-Max on coding and agentic tasks. No pricing, API availability, inference hardware requirements or weight licence are given in the visible text. The source text is cut off partway through the pretraining section, so anything after that point is not covered here.
“Where frontier open models like Kimi K3 remain ahead on raw capability, Beam's advantage is efficiency at inference time.”
— Reflection, Introducing Beam