AI NewsModels & agentsAnnouncement
Reflection unveils Beam, a 501B-parameter open-weight MoE with weights due later this month
Reflection says Beam is a 501B sparse Mixture-of-Experts model with 23B active parameters, pretrained on 23.8 trillion tokens and RL-trained on 10.5K NVIDIA GB300 GPUs over four weeks, with weights, a technical report and a model card promised later this month.

Image: Reflection
Why it mattersA team picking an open-weight model for coding and agent work gets a Western alternative to GLM, Qwen and DeepSeek that Reflection says runs at 3 to 4 times less inference compute for comparable reasoning scores.
A team picking an open-weight model for a coding or agent product has had to choose between GLM, Qwen and DeepSeek, all trained in China. Reflection, a Brooklyn startup founded in 2024, is offering a Western alternative, and the weights are not out yet.
Reflection published its "Introducing Beam" post on 5 October 2026. Beam is a sparse Mixture-of-Experts model with 501 billion total parameters and 23 billion active per token, built for coding, reasoning and agentic workloads. Reflection says it will release the weights, a technical report, a model card and developer artifacts later this month. Early access signups are open now, and the model is still being checked for safety and reliability before release.
The numbers Reflection reports
Reflection says Beam was pretrained on 23.8 trillion tokens from the web and licensed datasets. The reinforcement learning run that followed used 10.5 thousand NVIDIA GB300 GPUs for four weeks and produced over 100 million rollouts, with a maximum context length of 256,000 tokens. Reflection calls this one of the largest open RL runs to date. Training and grading used roughly 1.3 billion sandboxes, drawn from one million coding, agentic and STEM environments the team sourced.
On Reflection's own benchmark table, Beam scores 80.9 on SWE-Bench Verified, 65.5 on SWE-Bench Pro v1, 80.1 on Terminal-Bench 2.1, 97.8 on AIME 2026 and 90.5 on GPQA Diamond. Reflection says these put Beam on par with Z.ai's GLM-5.2 and approaching Qwen 3.8-Max on coding and agentic tasks, while Kimi K3 remains ahead on raw capability. None of these scores has been independently verified.
The efficiency claim is the one to check. Reflection says Beam matches GLM-5.2 on advanced reasoning benchmarks while using 3 to 4 times less inference compute, and that the difference grows against the 2-trillion-parameter class like Qwen 3.8-Max. Reflection's own figure caveats that the comparison is estimated, not measured: it counts active parameters times generated tokens, and excludes prompt prefill and serving overhead.
What a buyer can do with it today
The weights are not published, so nothing runs today. What is available is the benchmark table, the training recipe, and the signup page for early access. For a team that is costing a product on open-weight inference, Reflection has given enough detail to compare Beam against GLM-5.3, Qwen 3.8-Max, Kimi K3 and DeepSeek V4.1 Flash: a 23B active-parameter model has a different serving cost than a 2T dense model.
Reflection calls Beam a "workhorse" aimed at enterprise, public sector and sovereign-AI customers who want to train models on their own data in their own infrastructure. The company has raised about $4.7 billion from backers including Nvidia and Sequoia, per TechCrunch, and signed over $7 billion of chip deals with SpaceX and Nebius through 2029. The weights, the model card and a reproduction kit are what will decide whether the benchmark numbers are confirmed by anyone outside Reflection.
Source
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.