model release

Reflection AI unveils Beam: 501B-parameter open-weight MoE with 1M-token context

TL;DR

Reflection AI has unveiled Beam, a text-only mixture-of-experts model with 501 billion total parameters, 23 billion active, and a 1 million token context window. The company claims it matches Z.ai's GLM-5.2 on advanced reasoning benchmarks while using 3-4x less inference compute. Weights and the full technical report are due later this month.

3 min read
0

Reflection AI has unveiled Beam, a text-only mixture-of-experts (MoE) model with 501 billion total parameters, 23 billion active parameters, and a 1 million token context window. The company announced it Monday, October 5, 2026. According to Reflection, Beam matches Z.ai's GLM-5.2 on advanced reasoning benchmarks while using "3-4x less inference compute." Those performance claims have not been independently verified.

Specifications

Spec Beam
Architecture Mixture-of-experts, text-only
Total parameters 501B
Active parameters 23B
Pre-training data 23.8T tokens
Context window 1,000,000 tokens
Post-training High-compute reinforcement learning
Pricing Not yet disclosed
Training cutoff Not yet disclosed
Weights Due this month, per Reflection

Reflection describes Beam as a "workhorse model" for enterprises, the public sector, and developers. It says the model is trained for reasoning, coding, and agentic tasks at "a fraction of the token cost and inference time compute" of rivals.

For comparison, GLM-5.2 has roughly 744B total and 40B active parameters. Beam has about 33% fewer total parameters and about 42% fewer active parameters.

Performance claims

Reflection says Beam scores on par with GLM-5.2 on advanced reasoning benchmarks and outperforms today's leading Western open models. The company has not disclosed specific benchmark figures in the material available to us, and none have been independently confirmed.

Reflection's own results also show Beam outscoring Inkling, the open model from Thinking Machines Lab released in July, on four coding tests where both report scores. The comparison is not like-for-like: Inkling is multimodal, and Beam handles text only.

Availability

Reflection says it will release Beam's weights and full technical details this month. It also says the model will be distributed through hyperscalers and neoclouds, with integrations across open-source libraries at launch. Licensing terms have not been disclosed. Reflection did not respond to TechCrunch's requests for more information.

Company and compute

Reflection was founded in 2024 by two former Google DeepMind researchers and is based in Brooklyn. According to PitchBook, it has raised about $4.7 billion from backers including Nvidia, Sequoia Capital, and Lightspeed Venture Partners. Its last round valued the company at $25 billion pre-money.

This summer, Reflection signed deals collectively worth more than $7 billion with SpaceX and Nebius for access to Nvidia GB300 chips through 2029.

The company is targeting enterprises and sovereign customers with what it calls "AI factories." Institutions would train Reflection's models on their own proprietary data to build customized local systems. Reflection has begun testing a sovereign AI factory partnership with Shinsegae Group in South Korea. Axios reported that hedge funds and trading firms are also interested.

What this means

Beam is the strongest signal yet of a U.S. lab trying to compete with Chinese open-weight models on efficiency, not just capability. The headline claim is the 3-4x inference compute advantage. If it holds up, it matters more than any benchmark parity, because serving cost is the main variable for enterprises deploying open models at scale.

Three things need to be checked before the claims can be trusted. The first is the weights release, which will let outside teams reproduce the results. The second is the actual benchmark numbers, which Reflection has not yet put in front of independent evaluators. The third is the license, which determines whether sovereign and regulated buyers can use Beam the way Reflection's pitch implies.

The text-only design is a trade-off. It likely helps efficiency, but it leaves Beam without the multimodal capability that Inkling offers. The model's real test is whether it holds up on third-party reasoning and agentic evaluations once the weights are public.

Related Articles

model release

inclusionAI releases Ling 3.1 Flash: 560B MoE, 25B active, 262K context, free on OpenRouter

inclusionAI has released Ling 3.1 Flash, a hybrid reasoning mixture-of-experts model with 560B total and 25B active parameters and a 262K-token context window. It is listed as free on OpenRouter through NovitaAI. No benchmark scores have been published on the listing.

model release

Unbiased releases Pareto 26.10 Preview: 1M context, $0.80/$3.20 per 1M tokens on OpenRouter

Unbiased has listed Pareto 26.10 Preview on OpenRouter, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. The company says it targets research, coding, and agentic workflows, and warns the preview may change without notice. No benchmark scores have been published.

model release

Reka AI releases Rho-1, a 19B-parameter omni-model for text, image, video and robot control

Reka AI has released a research preview of Rho-1, a 19-billion-parameter omni-model that processes and generates text, images, video, and robot control actions in a single network. Reka says it uses no tool calls or external models. Context window, pricing, and benchmark scores have not been disclosed.

model release

China Telecom's Xing4.0-29B-A4B: 29B MoE, 4B Active, 256K Context, Trained Fully on Ascend NPUs

China Telecom AI's Xing4.0-29B-A4B (formerly the TeleChat line) is a mixture-of-experts model with 29B total and 4B active parameters and a native 256K context window, extensible to 512K. The company claims it is the first model of this scale trained entirely on Ascend NPUs with MindSpore. Community GGUF quantizations from Venastine-Research are already available.

Comments

Loading...