Reflection announces Beam, a 501B-parameter open-weight coding model with 23B active parameters
Reflection has announced Beam, its first open-weight model, a 501B-parameter mixture-of-experts system with 23B active parameters per token, built for coding, reasoning and agentic tasks. The company claims it matches GLM 5.2 on demanding reasoning tasks with three to four times less compute. Weights are due under Apache 2.0 later this month.
Reflection has announced Beam, a 501-billion-parameter mixture-of-experts model that activates 23 billion parameters per token. It is the company's first open-weight release. Model weights, a technical report and developer documentation are due "later this month" under the Apache 2.0 license. For now, only select users have access to an early version while final safety testing continues.
Beam is text-only and targets coding, logical reasoning and agentic workflows. Reflection says it can handle other media if the content is represented as text.
Specs and benchmarks
- Architecture: Mixture-of-experts, 501B total parameters, 23B active per token
- License: Apache 2.0 (weights only; Reflection keeps training data and pipelines proprietary)
- Context window: not yet disclosed
- Pricing: not yet disclosed
- Training cutoff: not yet disclosed
- Reasoning control: a tunable parameter trades response speed against reasoning depth
All figures below come from Reflection's own comparison and have not been independently verified.
| Benchmark | Beam | GLM 5.2 | GLM 5.3 | Kimi K3 | Qwen 3.8 Max | DeepSeek V4.1 Flash |
|---|---|---|---|---|---|---|
| Terminal Bench v2.1 | 80.1 | 81.0 | 88.2 | 88.3 | 86.6 | 90.6 |
| SWE Bench Pro v1 | 65.5 | 62.1 | NR | NR | 67.7 | NR |
| SWE Bench Pro v2-Hard | 77.2 | NR | 84.3 | 88.2 | NR | NR |
| DeepSWE v1.1 | 44.4 | 44.0 | 61.0 | 68.0 | 51.0 | 74.2 |
Beam also scores 80.9 on SWE-Bench Verified, 78.0 on SWE-Bench Multilingual and 34.6 on SWE Atlas Codebase QnA. Reflection's chart lists several rivals as "NR" (not reported) on individual tests. It also includes Inkling and Nemotron 3 Ultra, which Beam leads on the tests where both have scores.
Reflection claims Beam matches GLM 5.2 on demanding reasoning tasks while using three to four times less compute, and comes close to Qwen3.8-Max on coding and agent benchmarks. The company concedes that Kimi K3 still outperforms Beam on raw performance. The table also shows Beam trailing DeepSeek V4.1 Flash on both Terminal Bench v2.1 and DeepSWE v1.1.
Training
Beam combines large-scale pretraining with a compute-heavy reinforcement learning phase. Reflection says the RL phase ran on 10,500 Nvidia GB300 GPUs for more than four weeks and called it one of the largest training runs by an open lab. According to the company, benchmark scores kept rising through more than 80 million rollouts without plateauing.
Reflection also reports an unplanned gain in web browsing. The RL mix covered reasoning, software engineering and terminal tasks, but no browsing tasks. With web access, the model reportedly learned to query other language models and retrieve external documents. Demos included a live-updating New York City subway map, a small 3D game and a notebook for fine-tuning another model.
For safety, Reflection trained a separate model on behavioral guidelines, from hard prohibitions to standards such as factual accuracy and admitting uncertainty, and merged it with Beam. The company plans to publish safety results and open-source its evaluation methods.
Company context
Reflection was founded in 2024 by ex-Google DeepMind researchers Misha Laskin and Ioannis Antonoglou. It raised $130 million in seed funding in March 2025 and $2 billion at an $8 billion valuation in October 2025, with Nvidia among the investors. It has since signed billion-dollar compute deals with SpaceX and Nebius. Reflection says it is already training a successor aimed at closing the gap to the top open models.
What this means
Beam's pitch is cost per unit of capability, not a leaderboard lead. By Reflection's own numbers it sits roughly level with GLM 5.2 and clearly behind GLM 5.3, Kimi K3 and DeepSeek V4.1 Flash on several coding benchmarks. The 3-4x compute-efficiency claim is the figure to test once weights are public, since it determines whether Beam is cheaper to serve at equivalent quality. The headline claim of being the strongest open-weight model built outside China holds only against a narrow comparison set and rests on vendor-reported results. An Apache 2.0 license on a 501B-parameter model is notable for enterprises that want a Western-origin option without usage restrictions. Independent evaluations, the technical report and the final context window and serving requirements will decide whether it competes in practice.
Related Articles
Reflection AI unveils Beam: 501B-parameter open-weight MoE with 1M-token context
Reflection AI has unveiled Beam, a text-only mixture-of-experts model with 501 billion total parameters, 23 billion active, and a 1 million token context window. The company claims it matches Z.ai's GLM-5.2 on advanced reasoning benchmarks while using 3-4x less inference compute. Weights and the full technical report are due later this month.
inclusionAI releases Ling 3.1 Flash: 560B MoE, 25B active, 262K context, free on OpenRouter
inclusionAI has released Ling 3.1 Flash, a hybrid reasoning mixture-of-experts model with 560B total and 25B active parameters and a 262K-token context window. It is listed as free on OpenRouter through NovitaAI. No benchmark scores have been published on the listing.
TII's Falcon-Emirati-7B scores 84.83% on Alyah, a new Emirati-dialect Arabic benchmark
The Technology Innovation Institute (TII) released Falcon-Emirati-7B, a 7B-parameter model specialized for Emirati Arabic and built on Falcon-H1-Arabic. TII claims it scores 84.83% on the Alyah benchmark, ahead of every Arabic and multilingual model it compared against.
Reka AI releases Rho-1, a 19B-parameter omni-model for text, image, video and robot control
Reka AI has released a research preview of Rho-1, a 19-billion-parameter omni-model that processes and generates text, images, video, and robot control actions in a single network. Reka says it uses no tool calls or external models. Context window, pricing, and benchmark scores have not been disclosed.
Comments
Loading...