On October 5, 2026, Reflection AI announced Beam, its first open-weight model, designed for coding, reasoning, and agentic workloads. The company reported 501 billion total parameters, 23 billion active per inference, and benchmark results of its own; it planned to release the weights and technical materials later in October. NeoTeo previously covered Reflection AI’s plans for its first open-weight model; Reflection has now identified Beam as that model.
What Reflection announced
“Open-weight” refers to a model whose trained parameters—the values that shape its outputs—are made available for others to use or adapt. Reflection said it planned to release Beam’s weights under the Apache 2.0 license, along with a technical report, model card, and developer materials, later in October 2026. The announcement described Beam as undergoing final red-teaming and evaluations, with early access offered to a select group.
Reflection describes Beam as a text-only language model built with a sparse mixture-of-experts architecture. In this design, the model has specialized components, or “experts,” and uses a subset of them for a given inference. The company says Beam has 501 billion parameters in total and 23 billion active per inference. Those figures describe different things: the first counts the full model, while the second is the active-parameter count Reflection reports for each inference.
Beam’s announced specifications and training
Reflection says it pretrained Beam on 23.8 trillion tokens drawn from web and public sources, as well as proprietary licensed datasets. It reports an effective context length of 1 million tokens after midtraining. Context length is the amount of text a model can take into account in a task.
That figure is separate from the 256K-token maximum context used in the reinforcement-learning run described by Reflection. The company says that run used 10,500 NVIDIA GB300 GPUs over four weeks and generated more than 100 million rollouts—training attempts in which the model works through a task and receives feedback.
What Reflection says about benchmarks and compute
Reflection’s published benchmark table gives Beam a score of 80.9 on SWE-Bench Verified and 80.1 on Terminal-Bench 2.1. It also reports 36.2 on Humanity’s Last Exam without tools. These are results Reflection reported for Beam.
The company says Beam reached scores comparable to GLM-5.2 on advanced reasoning benchmarks with 3–4× less inference compute. Reflection’s estimate covers generation forward-pass operations; it excludes prompt prefill, context-dependent attention operations, and serving overhead. It therefore describes an approximate inference-compute comparison, not a full measure of deployment cost.
Reflection said it was still conducting final red-teaming and evaluations when it announced Beam, and planned to publish the weights, model card, technical report, and developer materials later in October 2026.