Reflection AI Beam Debuts With 501B Parameters
-
- by THEFLGHT,
- October 06, 2026
- in Artificial-Intelligence
Reflection AI Beam debuted on October 5, 2026, as the company’s first open-weight model, with a 501-billion-parameter design aimed at coding, reasoning and AI agents. The Nvidia-backed company says only 23 billion parameters are active for a given token, a choice intended to reduce the compute needed to run a very large model.
The announcement is an early-access debut, not a public release of the weights. Reflection says Beam remains in final safety testing and evaluation; it plans to publish the weights, technical report, model card and developer materials later in October.
Reflection’s Beam announcement establishes four key facts:
- 501 billion total parameters, with 23 billion active.
- A focus on coding, reasoning and agent tasks.
- More than 100 million training rollouts on Nvidia GB300 GPUs.
- Early access now, with public weights promised later this month.
Reflection AI Beam Opens in Early Access
Reflection’s announcement describes Beam as a sparse mixture-of-experts model. Instead of running every part of its 501-billion-parameter network for each token, the model routes work through a smaller selection of experts. Its 23-billion active-parameter figure describes that selected portion, not the model’s total size or memory footprint.
That distinction matters for buyers weighing capability against serving cost. Fewer active parameters can lower the arithmetic needed to generate a token, but total deployment cost also depends on memory, batch size, response length, infrastructure and the software wrapped around the model.
Interested developers can seek early access while Reflection completes red-teaming and evaluations. The company has not yet released the downloadable weights or the full technical and safety documentation, so claims about independent deployment and customization remain prospective until those materials arrive.
Reuters reported that Reflection is positioning Beam against lower-cost Chinese open-weight models used for code and agentic tasks. Founded by former DeepMind researchers Misha Laskin and Ioannis Antonoglou, the company is entering a crowded field where developers can compare downloadable models on capability, operating expense and control.
Beam Benchmarks Put Coding Ahead of Some Rivals
Reflection reports an 80.9% score for Beam on SWE-bench Verified, 77.2% on SWE Bench Pro v2-Hard and 80.1% on Terminal Bench v2.1. It also lists 78.7% on MCP Atlas, which tests tool use, and 37.0% on the public AutomationBench split.
The table shows clear limits alongside the strong results. On the same reported DeepSWE v1.1 comparison, Beam scores 44.4 against 61.0 for GLM 5.3, 68.0 for Kimi K3 and 74.2 for DeepSeek V4.1 Flash. Reflection itself says Kimi K3 remains ahead on raw capability.
These figures come from Reflection’s published comparison, not a complete independent evaluation of a downloadable Beam release. Different harnesses, tool budgets and model settings can change benchmark results, while several rivals have unreported scores on particular tests. Developers will need to reproduce task-specific performance when access broadens.
Reflection argues that Beam’s selling point is capability per unit of inference compute. It estimates comparable advanced-reasoning performance to GLM-5.2 with three to four times less generation compute. The estimate excludes prompt processing, attention and serving overhead, so it does not establish a measured reduction in a customer’s total bill.
How Reflection Trained the 501B Model
Reflection says it pretrained Beam on 23.8 trillion curated tokens drawn from web material, public sources and proprietary licensed data. Its subsequent reinforcement-learning run used 10,500 Nvidia GB300 GPUs over four weeks and produced more than 100 million rollouts, the attempted task trajectories used to improve the model.
The company describes a large training environment behind those numbers: roughly one million coding, agent and STEM tasks, around 1.3 billion sandbox uses for training and grading, and an average of 110,000 concurrent rollouts. Those are company-reported infrastructure figures rather than independently audited measures.
To keep that work moving, Reflection built an asynchronous system that can generate task attempts while training continues. The company says model updates reached its inference fleet in a median of about 12 seconds, and that it recovered from 71 inference incidents without ending the overall training run.
Beam also exposes a reasoning-effort setting intended to trade shorter responses for additional work on harder tasks. Reflection says its training rewarded successful solutions while discouraging unnecessary tokens, then allowed longer reasoning when it improved performance. The practical cost and quality effects will depend on the final product configuration.
Weights and Independent Testing Are the Next Tests
Open weights could let organizations run and adapt Beam under their own controls, but the conditions of that use will depend on the forthcoming release. The model card and technical report should clarify license terms, safety findings, hardware requirements and how the published evaluations were conducted.
For developers, Beam’s debut adds another potential option between closed frontier APIs and existing open-weight systems. Its sparse design may be attractive where coding and tool use dominate, provided real deployments confirm the advertised performance and deliver acceptable latency and cost.
The immediate milestone is Reflection’s promised October release of weights and documentation. Until then, Beam is a significant model announcement with substantial disclosed training detail, but its strongest claims remain company-reported and access is limited.
0 Comments:
Leave a Reply