QUICK ANSWER

Reflection has announced a 501-billion-parameter mixture-of-experts model with 23 billion parameters active per token. That smaller active count can reduce computation, but it does not mean a 23B checkpoint: raw BF16 weights are about 1.0 TB before runtime overhead. Beam is currently waitlist-only; weights are planned later in October.

What Reflection announced

Reflection introduced Beam on October 5 as its first open-weight model, aimed at coding, reasoning and agentic workloads. The company describes a sparse mixture-of-experts (MoE) model with 501 billion total parameters and 23 billion active for each token. As of October 6, Reflection says selected users can request early access; it plans to publish the weights, model card, technical report and developer tools later this month under Apache 2.0. That means readers cannot yet inspect a public checkpoint or verify its final runtime requirements.

Why 23B active does not mean a 23B model download

An MoE model routes each token through some of its expert networks rather than computing through every expert at once. This can lower the computation used for a token compared with a dense model of the same total size. The checkpoint still contains the experts: a 2025 ACL paper on MoE compression explains that fewer active parameters do not reduce total parameter count, and that expert weights create substantial memory pressure. Runtime strategies can shard weights across GPUs or offload them to system memory, but those choices need supported software and trade off speed, complexity or both.

How much memory could Beam’s weights take?

A simple lower-bound estimate multiplies parameter count by bytes per weight. For 501 billion parameters, BF16 at two bytes per weight works out to about 1,002 GB (roughly 933 GiB) for the weights alone. FP8 at one byte is about 501 GB; four-bit storage at half a byte is about 251 GB before scales and metadata. These are arithmetic estimates, not a tested Beam build: Reflection has not yet published the weight files or documented supported precisions. Real inference also needs memory for the runtime, activations, cache and other buffers.

Illustrative storage formatRaw weight estimate for 501B parametersWhat this does—and does not—tell you
BF16 · 2 bytes/parameterAbout 1,002 GBA raw-weight lower bound; not a complete serving requirement.
FP8 · about 1 byte/parameterAbout 501 GBA hypothetical estimate; Beam’s supported format is not published.
4-bit · about 0.5 byte/parameterAbout 251 GBA hypothetical estimate before quantization metadata and runtime memory.

Can one consumer GPU run Reflection Beam?

Not as a conventional, fully resident BF16 model on one card. NVIDIA lists 32 GB of memory for the GeForce RTX 5090 and 141 GB for the data-center H200; both are below even the rough 4-bit weight-only estimate above. A multi-GPU server, aggressive offload or a future compressed checkpoint could change the deployment options, but the launch post does not specify a supported inference configuration. The reported 6,144-GB300 pretraining cluster and 10,500-GPU reinforcement-learning run describe Reflection’s training work, not the hardware a user needs to serve Beam.

What do Beam’s benchmark numbers show?

Reflection’s launch post reports Beam at 80.1 on Terminal-Bench v2.1. Its table lists Kimi K3 at 88.3 and Qwen 3.8 Max at 86.6 on the same benchmark, while Beam leads some other entries or sits close to them. Those are company-published results, not an independent evaluation of Beam. Reflection also presents estimated inference-compute comparisons and explicitly excludes prompt-prefill, context-dependent attention and serving overhead from that estimate. Treat the table as an early vendor report: the model card, harness settings, model versions and independent reruns will matter before calling a winner.

What local AI builders should do now

Do not buy a GPU based on the 23B active number alone. Wait for Reflection’s checkpoint and model card, then check the actual quantized files, supported inference engines, memory measurements at your intended context length, license text and download source. If Beam is available only through a hosted endpoint at first, compare its live price and usage limits before planning a local build. For hardware choices you can make today, start with your actual workload and current software support in our AI workstation planner, and use our local LLM VRAM guide to understand the memory trade-offs.

CLEAR ANSWERS

Frequently asked questions

What is Reflection Beam?

Beam is Reflection’s announced open-weight language model for coding, reasoning and agentic tasks. Reflection describes it as a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active per token.

Why does Beam have 501B total parameters but 23B active?

The model routes each token through a subset of its expert networks, so fewer parameters are used for that token’s computation. The total checkpoint still includes the experts, so active parameters do not equal the full model’s storage requirement.

Can I run Reflection Beam on an RTX 5090?

There is no published supported build yet. The RTX 5090 has 32 GB of memory, while simple weight-only arithmetic estimates about 251 GB even at four bits for 501 billion parameters, before metadata and runtime memory. Wait for Reflection’s model card and supported quantized files before planning a deployment.

Can I download Beam now?

Reflection says selected users can request early access. The public weights, model card, technical report and developer artifacts are planned for later in October 2026; they were not available in the October 5 announcement.

Has Beam independently beaten Kimi K3 or Qwen 3.8 Max?

The October 5 announcement provides Reflection’s own benchmark table, not an independent evaluation. In its Terminal-Bench v2.1 row, Reflection lists Beam at 80.1, Kimi K3 at 88.3 and Qwen 3.8 Max at 86.6. Independent tests with published settings are still needed.

Does Reflection publish Beam API pricing?

The launch announcement does not list a public API price. Check Reflection’s current platform information when early access or public distribution opens.

FOLLOW THE SOURCE

Sources & further reading

Primary sources checked Oct 6, 2026. Vendor statements are attributed; editorial advice is our own.

  1. 1
  2. 2
  3. 3
  4. 4