Skip to main content

Reflection AI Opens Beam, a 501B-Parameter Coding MoE Under Apache 2.0

Reflection AI announced Beam, an open-weight 501-billion-parameter Mixture-of-Experts model with 23 billion active per token, scoring 77.2 on SWE-Bench Pro v2-Hard. Weights ship under Apache 2.0 later this month with a technical report. The company positions Beam as a Western workhorse for coding and agentic workloads, claiming three to four times less inference compute than larger rivals.

On this page

A 501B MoE trained in under four weeks

Reflection AI, the Nvidia-backed startup founded by ex-DeepMind researchers to be the Western answer to DeepSeek's open-weight strategy, announced Beam on 5 October: a sparse Mixture-of-Experts with 501 billion total parameters and 23 billion active per token, trained on 23.8 trillion tokens. The training claims are the quiet headline. Pretraining ran on 6,144 GB300 GPUs in under four weeks at 92.3 percent goodput, and the reinforcement-learning stage used 10,500 more GPUs, over 100 million rollouts, and about 1.3 billion sandboxes with up to 170,000 running concurrently, which the company calls one of the largest RL runs by any open lab to date.

Apache 2.0 weights, a technical report, and a model card are due later this month, with an early-access waitlist on the company's platform in the meantime. The announcement also commits to open-source safety evaluations, extending the argument from the company's homepage that open models let researchers probe risks that closed development bottlenecks.

The benchmarks, honestly framed

Beam scores 77.2 on SWE-Bench Pro v2-Hard, 80.9 on SWE-Bench Verified, 97.8 on AIME 2026, and 90.5 on GPQA Diamond. The company's own comparison table is refreshingly direct about where those numbers sit: GLM 5.3 leads the same benchmark at 84.3 and Kimi K3 at 88.2, meaning Beam is not claiming the open-weight crown on raw capability. The claimed edge is efficiency, matching or beating GLM 5.2's quality while using three to four times less inference compute, and nearing much larger models on coding and agentic tasks.

That positioning makes Beam an enterprise workhorse pitch rather than a leaderboard swing: deploy it on your own infrastructure, air-gapped if needed, tune it with the published recipes, and accept that a 2-trillion-parameter frontier model still outperforms it when raw capability is the requirement. A Hacker News thread on the announcement drew over 400 points, reflecting how much demand exists for exactly this tier of open model.

The Western open-weight bet

Beam is the first major release from the strategy Reflection has funded openly: two billion dollars raised at an eight billion valuation in 2025, explicitly to build the open-weight counterpart to DeepSeek, and reports this year of a round approaching 25 billion. Whether Beam justifies that trajectory will show up in what ships this month, because the announcement's numbers are the company's own and the weights are the falsifiable part.

For the local and self-hosted community, the calculus is straightforward. A 501B sparse model with 23B active sits in the same deployment class as other large open MoEs: multi-GPU servers and expert-caching engines rather than gaming laptops, but with permissive licensing, published containers and agentic harnesses, and a vendor whose business depends on people running the weights themselves. The open-weight field now has credible Western entries at the frontier-adjacent tier, and procurement decisions that defaulted to a single Chinese option have alternatives with different governance attached.

CuriousLM runs supported AI models locally on your device. Try CuriousLM.