Reflection Beam: 501B open-weight model, weights later in October 2026

Reflection Beam is a 501B sparse MoE with 23B active parameters for coding and agent tasks; Reflection plans Apache 2.0 weights later in October 2026.

Reflection Beam is a sparse Mixture-of-Experts model with 501 billion total parameters and 23 billion active parameters per token. Reflection announced it on 2026-10-05 for coding, reasoning, and agentic tasks, and plans to release the weights under Apache 2.0 later in October 2026 while final red-teaming and evaluations are underway. As of 2026-10-06, access is through a waitlist.

Key facts about Reflection Beam

  • Beam has 501 billion total parameters and 23 billion active parameters per token.
  • Reflection reports a 1 million-token effective context length, and its architecture has 52 layers.
  • Reflection says Beam is comparable to GLM-5.2 on advanced reasoning benchmarks while using 3 to 4 times less estimated inference compute.
  • Reflection pretrained Beam on 23.8 trillion tokens in under four weeks on 6,144 NVIDIA GB300 NVL72 GPUs.
  • Its reinforcement-learning run used 10,500 NVIDIA GB300 GPUs for four weeks, generated more than 100 million rollouts, and used about 1.3 billion sandboxes for training and grading.
  • The planned release includes the weights, technical report, model card, and developer artifacts under Apache 2.0.

Reflection Beam's architecture and context window

Beam is text-only and uses sparse expert routing. Each token activates 23 billion of the model's 501 billion parameters. The architecture uses interleaved local and global attention, fine-grained routed experts, auxiliary-loss-free load balancing, depth-based residual scaling, SandwichNorm, elementwise attention gating, and FP32 residual accumulation.

Reflection reports that the busiest expert's load averaged 1.04 times the mean at the end of pretraining. The official post also reports bounded residual-stream norms across all 52 layers. Midtraining extended Beam's effective context length to 1 million tokens.

Beam exposes a reasoning-effort parameter. Lower settings favor shorter responses. Higher settings allow longer reasoning on demanding tasks, so a developer can trade token use against performance.

Reflection Beam's benchmark and efficiency claims

The benchmark figures below come from Reflection's launch post. TechCrunch and The Next Web both note that the company's performance claims had not been independently verified as of 2026-10-06.

Benchmark

Beam

Other scores in Reflection's table

Terminal Bench v2.1

80.1

GLM-5.2: 81.0; GLM-5.3: 88.2; Kimi K3: 88.3; Qwen 3.8 Max: 86.6

SWE-bench Verified

80.9

Inkling: 77.6

SWE Bench Pro v2-Hard

77.2

GLM-5.3: 84.3; Kimi K3: 88.2

Reflection's 3 to 4 times efficiency claim is a theoretical forward-pass comparison, not a measurement of production serving cost. Its formula is:

FLOPs ≈ 2 × active parameter count × mean generated tokens per attempt

The estimate counts reasoning and answer tokens, uses active rather than total parameters for MoE models, and excludes prompt prefill, context-dependent attention operations, and serving overhead. That makes it useful for comparing the published estimate, but not for predicting the cost or latency of a deployed Beam service.

Reflection Beam's training scale

Reflection says Beam's pretraining used 23.8 trillion tokens from web, public, and proprietary licensed datasets. The run finished in under four weeks on 6,144 NVIDIA GB300 NVL72 GPUs. Reflection reports 92.3% goodput and nine semi-automatic rewinds attributed to gradient-norm spikes or suspected silent data corruption.

The reinforcement-learning campaign ran for four weeks on 10,500 NVIDIA GB300 GPUs. It generated more than 100 million rollouts with a maximum context length of 256,000 tokens, and training and grading used approximately 1.3 billion sandboxes. Reflection also reports one million coding, agentic, and STEM environments, an average of 110,000 concurrent rollouts, and up to 170,000 concurrent sandboxes.

Beam used asynchronous policy gradients. Reflection says learning stayed numerically stable when samples were generated up to 107 weight versions, or about one day, behind the current policy. Safety alignment combined a large-scale reinforcement-learning teacher with a dedicated safety teacher through multi-teacher on-policy distillation.

Reflection Beam's release status

Reflection's launch post on X appeared at 19:11 UTC on 2026-10-05. The official announcement says Beam was still undergoing final red-teaming and evaluations, with early access offered through a waitlist at platform.reflection.ai.

Reflection plans to release the weights in October 2026 under Apache 2.0, together with documentation and the full stack for running, evaluating, and fine-tuning the model. The release is planned to include integrations with open-source libraries and harnesses. Until that release happens, the public announcement offers waitlist access rather than a public weight package.

Sources

Last verified: 2026-10-06.

Spotted an outdated or wrong claim? Agents can report it with evidence throughPOST /api/feedback; an editor checks every report. See llms.txt for the agent API.