LocalsOnlyevaluationsHow to read this
← Models

Reflection Beam

Published frontier peer. Announced open-weight, weights pending: early-access waitlist only, not a closed API and not downloadable yet. Reflection says weights, tech report, and model card ship later this month under Apache 2.0. Sparse MoE, 501B total / 23B active, 52 layers, interleaved local and global attention, 1M-token effective context, text-only, with a reasoning-effort parameter. Vendor-claimed launch-table cites, not independently verified. The post does not state harness/agent, reasoning effort, or sample counts. Banked depth rows: AutomationBench public 37.0 (version not stated; filed as automationbench-public; protocol differs from our desk runs), DeepSWE v1.1 44.4 (not an Index column), IFBench 79.7 (not an Index column). GPQA Diamond 90.5 is banked as a published cite only (reasoning effort not stated); it is not on the GPQA ranking. Index Terminal-Bench 2.1 and SWE-bench Verified stay empty: the table lists 80.1 and 80.9, and those columns require a disclosed harness/agent before a value sits in the cell. Local fit is pending weights and untested. By size alone, 501B does not fit one GB10 (128 GB) at 2-bit or above; a ~3-bit pack might fit 2× GB10 (256 GB) with little headroom. No weights or community quants exist yet, so there is no EXAMPLE desk row. Not a LocalsOnly desk run. https://reflection.ai/blog/introducing-beam