Reflection AI's Beam Challenges DeepSeek With 501B Open-Weight Reasoning Model

Reflection AI's first open-weight model is a 501B Mixture-of-Experts system that reasons 3-4x more efficiently than comparable open models.

·
·
·
Reflection AI's Beam Challenges DeepSeek With 501B Open-Weight Reasoning Model
Read7 min
TypeNews
  • Reflection AI announced Beam, a 501B parameter MoE model with 23B active, Apache 2.0 licensed
  • Claims 3-4x less inference compute than GLM-5.2 at comparable reasoning scores
  • Trained with 100M+ RL rollouts on 10.5K Nvidia GB300 GPUs over 4 weeks
  • Pretrained on 23.8T tokens; extended to 1M token context during midtraining
  • Weights, technical report, FP8 and NVFP4 quantized versions releasing this month
  • Early access via the Reflection platform waitlist

Reflection AI unveils Beam, a 501B open-weight model built for efficient reasoning

Reflection AI has introduced Beam, a sparse Mixture-of-Experts model for coding, reasoning, and agentic workloads. The New York startup positions it as a Western counterpart to Chinese open-weight systems from DeepSeek, Kimi, and GLM. Beam contains 501 billion parameters but activates about 23 billion for each token, reducing inference computation while retaining a much larger pool of learned weights.

Reflection says it will release the weights, technical report, model card, and developer artifacts later this month. Planned downloads include FP8 and NVIDIA’s NVFP4 quantizations. Open-weight here means the model parameters will be available for commercial use subject to the license terms; the proprietary training data and full training pipeline are not part of the announced release.

The 23-billion active-parameter count reduces arithmetic per token, but self-hosting still requires access to the full expert set. The raw weights would occupy roughly 501 GB in FP8 or 251 GB in 4-bit form before accounting for quantization metadata, the key-value cache, and runtime overhead. Most production deployments will therefore require multiple accelerators, and support for FP8 and NVFP4 varies by GPU and inference engine.

Beam’s benchmark tradeoff

Reflection reports that Beam uses one-quarter to one-third as much inference compute as comparable open models on selected reasoning tasks. Its evaluations place the model near GLM-5.2 while consuming fewer generated tokens and floating-point operations, and within range of Alibaba’s larger Qwen 3.8-Max. Kimi K3 and DeepSeek V4.1 Flash retain higher scores on several coding benchmarks.

Reflection’s reported Beam benchmark results
Benchmark Reported score What it measures
SWE-bench Verified 80.9 Resolution of validated GitHub issues
Terminal-Bench v2.1 80.1 Multi-step work in terminal environments
SWE-bench Pro v1 65.5 Complex repository-level software tasks
SWE-bench Multilingual 78.0 Software engineering across programming languages
DeepSWE v1.1 44.4 Long-horizon software engineering tasks

On DeepSWE, the text-only version of Humanity’s Last Exam, and Terminal-Bench 2.1, Reflection’s plots place Beam on a stronger score-to-compute frontier than similarly sized open models. In practical terms, the company reports either a higher score for the same inference budget or comparable performance with fewer generated tokens and FLOPs.

These results come from Reflection’s own evaluation harness. Prompt templates, tool access, token budgets, sampling settings, and sandbox configuration can materially change agent benchmarks, so independent reproduction will be necessary once the weights and serving code are available.

Scatter plots comparing Beam’s reasoning efficiency with GLM-5.2, Nemotron 3 Ultra, and Qwen models
Reflection compares benchmark scores with generated tokens and inference FLOPs.

RL at 100 million rollouts

Beam’s high-compute reinforcement-learning phase generated more than 100 million rollouts over four weeks on 10,500 NVIDIA GB300 GPUs. A rollout is one attempted solution or sequence of interactions from which the training system derives a reward. Reflection describes the run as the largest publicly documented RL campaign for an open-weight model, compared with 30 million rollouts for its earlier Inkling model and 753,000 for Xiaomi’s MiMo.

Reflection reports continued gains as it increased RL compute, without a plateau during the run. The published scaling plots cover DeepSWE, Humanity’s Last Exam, and Terminal-Bench, so the claim does not establish that every capability would continue improving at the same rate.

The training system used asynchronous policy-gradient updates to keep generation and optimization running in parallel. Asynchronous RL introduces staleness because a rollout may have been generated by an older checkpoint by the time training consumes it. Reflection says Beam remained stable while learning from trajectories more than a day old and generated by policies as many as 107 weight versions behind the current checkpoint.

Systems figures from the run

  • Concurrent work: An average of 110,000 rollouts remained active during training.
  • Sandbox capacity: The system ran as many as 170,000 sandboxes concurrently, with 90% ready within 10 seconds.
  • Weight distribution: New checkpoints reached the inference fleet in a median of about 12 seconds through a hierarchy using RoCE, or RDMA over Converged Ethernet, and NVIDIA NVLink.
  • Environment volume: Training and grading created approximately 1.3 billion sandbox instances.
  • Task diversity: Reflection sourced one million high-quality coding, agentic, and STEM environments for the RL campaign.
  • Pretraining utilization: A separate cluster of 6,144 GB300 NVL72 GPUs reached 92.3% goodput, the share of wall-clock capacity spent on useful training work.
Charts showing Beam benchmark scores rising with additional reinforcement-learning rollouts
Reflection reports continued benchmark gains as the number of RL rollouts increased.

Rewarding concise solutions

Reflection trained Beam with a controllable length penalty that rewards successful solutions while discouraging unnecessary tokens. During early RL, benchmark scores rose as completion lengths fell, indicating that the model was learning shorter successful trajectories. Completion lengths later increased as Beam acquired more multi-step and tool-using behavior, accompanied by further score gains.

The inference interface exposes a reasoning_effort parameter for adjusting the tradeoff among latency, token cost, and task accuracy. Lower settings target routine work with shorter outputs, while higher settings allocate more inference compute to difficult tasks. Production evaluations will need to measure this control against each workload because the optimal setting depends on task complexity and serving constraints.

In a transfer experiment limited to reasoning, software engineering, and terminal tasks, Reflection observed gains on browsing benchmarks that were absent from that experiment’s RL mixture. With external tools available, Beam also learned to query other language models and call OCR APIs when reading documents. That behavior supports broader agent workflows while making tool permissions, network controls, logging, and protection against prompt injection relevant deployment requirements.

A 23.8-trillion-token base

Beam was pretrained on 23.8 trillion tokens drawn from filtered web content and licensed proprietary datasets. Reflection’s architecture combines local attention for nearby tokens, global attention for long-range information, and fine-grained routed experts that send each token through a small subset of the network.

The final base model achieved nearly uniform expert utilization, with the busiest expert handling 1.04 times the average load. Balanced routing keeps accelerator capacity productive and reduces the risk that lightly used experts receive too little training before reinforcement learning begins.

Reflection says its filters removed about 95% of raw internet tokens. Its classifiers also recovered roughly 1.8 trillion tokens that conventional open-source web filters would have discarded, including 87% of the company’s curated web-code data. Those figures describe Reflection’s internal filtering comparison and cannot be reproduced without the underlying datasets and classifiers.

A midtraining stage extended Beam’s context window to one million tokens before RL. Supporting that window at the architecture level does not establish uniform accuracy across a million-token prompt, and long contexts can sharply increase key-value-cache memory and latency. Repository retrieval, long-document recall, and tool-use tests will provide more useful deployment evidence than the maximum context figure alone.

Licensing, access, and deployment

Founded in 2024 by former DeepMind researchers, Reflection is targeting organizations that need to run models in their own cloud, on-premises, or in air-gapped environments. Apache 2.0 permits modification and commercial distribution subject to its notice and license conditions, giving teams more deployment control than hosted-only APIs provide.

The company’s disclosed infrastructure commitments help explain the scale of the training run. Reflection reported a $6.3 billion compute agreement with SpaceX for access to NVIDIA GB300 systems at the Colossus 2 data center, followed by a $1 billion agreement with cloud provider Nebius. It says a larger successor to Beam is already in training.

Early access is available through the Reflection platform before the weight release. Evaluations for coding and long-horizon agents should record task success, generated tokens, wall-clock latency, full-fleet memory use, tool-call reliability, and sandbox security. Those measurements will show whether Beam’s reported compute efficiency survives a team’s own prompts, repositories, tools, and serving stack.

Trending
  • No trending articles

Comments

avatar

Next Reads