Imagine you're a baseball scout, but instead of watching a prospect play a full nine-inning game, you show them film of a perfect at-bat and ask: could you have made that swing, given what the pitcher just threw? If they can recognize the right action in context — even without running the full game themselves — that tells you something real about their ceiling. This paper does exactly that for coding agents. The committed claim: you can predict how well a base language model will perform as a post-trained coding agent on SWE-bench Verified by measuring three cheap proxy signals — without ever letting the base model attempt the full multi-step agentic task. The core insight is that base models fail agentic coding benchmarks not because they lack the reasoning capability, but because they can't reliably produce well-formed tool invocations to even start the harness. Pass@K from a cold start is therefore a terrible predictor. Instead, the authors replay successful agent trajectories from post-trained models, identify the 'decisive step' — the first code-changing action whose cumulative patch flips tests from failing to passing — and then probe the base model's competence at that exact moment. The three screens are elegantly simple. Decisive-Action BPB (bits per byte) asks: how much probability mass does the base model assign to the known-correct action at the decisive step? Patch MCQ presents a multiple-choice test between the correct patch and alternatives rejected by the verifier. Prefix-conditioned pass@K lets the base model generate freely from the decisive context and checks whether any completion passes the test suite. All three avoid the cold-start tool-calling problem by conditioning on the trajectory prefix — the base model never has to drive the harness itself. The validation covers ten pairs of publicly available base and post-trained models, and the headline result is that all three screens rank the cohort in close agreement with post-trained SWE-bench Verified pass@1. This is a correlation claim, not a causation claim, and the authors are relatively careful about that distinction. The method requires only successful trajectories and a verifier — no access to the post-training recipe, reward model, or proprietary infrastructure. The architecture here sits squarely in the evaluation-as-inference family: you're using the base model as a scorer or conditional generator, not as an autonomous agent. The compute cost is dominated by forward passes through the base model at a single trajectory step, which is orders of magnitude cheaper than running full agentic loops. The method leans on having a corpus of successful post-trained trajectories, which creates a bootstrapping dependency — you need at least one good post-trained model to generate the signal you'll use to screen future base models. Integrity is reasonable but not airtight. The benchmark is SWE-bench Verified, a community standard for agentic coding. The ten model pairs are public. But the validation is purely correlational across a relatively small cohort, and there's no pre-registration or holdout temporal split. The screens are also evaluated post-hoc on the same benchmark whose trajectories define the decisive steps — a circularity the authors acknowledge implicitly but don't stress-test. The big unanswered question is whether the ranking holds when you move to a genuinely new benchmark with different task distributions. The practical value for large-scale LLM development is clear: if you're training dozens of base checkpoints and each full agentic post-training run costs six or seven figures in compute, a cheap screening pass that correctly triages which checkpoints to invest in saves real money. The paper doesn't quantify the compute savings explicitly, but the asymmetry between a few forward passes at a single trajectory step versus a full post-training pipeline is stark. Whether these screens generalize beyond SWE-bench Verified and beyond the current generation of models is the load-bearing question the paper doesn't answer.