Imagine you're training for a marathon, but instead of just running more miles, you hire a coach who watches you run a hundred million practice laps, correcting your form each time — and crucially, your coach keeps giving feedback even when they're watching yesterday's footage. That's the core mechanism behind Beam: Reflection AI's bet that massive-scale reinforcement learning, specifically asynchronous policy gradient methods with extreme staleness tolerance, can turn a well-pretrained base model into a frontier-competitive reasoning and coding agent at a fraction of the inference cost. Beam is a sparse Mixture-of-Experts model: 501 billion total parameters, but only 23 billion active per forward pass. It was pretrained on 23.8 trillion tokens and then subjected to what Reflection claims is one of the largest RL campaigns by any open lab — over 100 million rollouts generated on 10,500 NVIDIA GB300 GPUs over four weeks, consuming roughly 1.3 billion sandboxed execution environments. The central claim is not raw capability supremacy but inference efficiency: Beam reportedly matches GLM-5.2 on advanced reasoning benchmarks while using 3-4× less compute, and narrows the gap to larger models like Qwen 3.8-Max that sit in the 2T+ parameter family. The benchmark picture is mixed in exactly the way you'd expect from an honest disclosure. On SWE-Bench Verified, Beam scores 80.9 — strong, ahead of Inkling (77.6) and Nemotron 3 Ultra (70.7). On Terminal Bench v2.1, it hits 80.1 but trails GLM 5.3 (88.2), Kimi K3 (88.3), and DeepSeek V4.1 Flash (90.6). On DeepSWE v1.1, it scores 44.4 versus Kimi K3's 68.0 and DeepSeek's 74.2. The pattern is clear: Beam is competitive within its weight class and punches above it on efficiency, but frontier closed models and larger open models still lead on raw capability. The paper says this plainly, which is notable. The RL infrastructure story is technically the most interesting part. Asynchronous policy gradients at this scale create a severe staleness problem — rollouts generated by a model checkpoint from a day ago are training the current model. Reflection claims to have solved this with novel algorithms that maintain stable learning even when training on samples 10^7 weight versions stale. Their Figure 4 shows stable numerics under these conditions. If true, this is a genuinely useful infrastructure contribution, because staleness is the primary bottleneck for scaling asynchronous RL. The controllable length penalty is also worth noting: early in training, the model learned to solve tasks with fewer tokens before later expanding token usage to tackle harder problems. The generalization evidence is suggestive but thin. Reflection reports that training on reasoning, software engineering, and terminal tasks produced gains in browsing tasks that were never in the RL mixture — including the model spontaneously learning to query other LLMs and use OCR APIs. This is the kind of emergent capability claim that deserves independent verification. The demos (NYC subway dashboard, Gemma-4 fine-tuning notebook) are illustrative but not systematic evidence. The critical gap is that weights, technical report, model card, and developer artifacts are not yet released — they're promised 'later this month.' The model is still undergoing red-teaming. This means we're evaluating a press release, not a reproducible artifact. The benchmark numbers are self-reported, the comparison methodology uses estimated FLOPs rather than measured inference costs, and no independent lab has verified any claim. The sign-up-for-early-access model is standard for generating buzz, but it means the open-weight promise is currently just a promise. Reflection's previous model launch (Reflection 70B in 2024) was met with significant skepticism over benchmark reproducibility. That history makes independent verification of Beam's claims even more important. The technical narrative is plausible and well-constructed — the RL scaling curves, the staleness management, the efficiency framing are all internally consistent. But plausible is not verified. The real test comes when the weights drop and the community runs its own evals.