Imagine you're grading a student's math proof, and you suspect some steps are real reasoning while others are just pattern-matching off the specific wording of the question. So you rephrase the question — same math, different words — and watch which steps wobble. The steps that stay firm are load-bearing logic; the ones that shift are likely cribbing from surface features. That's exactly what SCAPO does to large language models during reinforcement learning. The committed claim: GRPO's uniform advantage assignment — giving every token in a correct response the same reward signal — actively reinforces spurious dependencies on task-irrelevant prompt features. SCAPO fixes this by measuring each token's probability drift under semifactual interventions (rephrased prompts that preserve the underlying problem) and downweighting tokens that are unstable. This is not a new RL algorithm; it's a causally motivated credit-assignment patch grafted onto GRPO's existing machinery. The results are concrete. On Qwen3-4B-Base, SCAPO beats GRPO on AIME 2024–2026 by 5.63 percentage points. On Qwen3-1.7B-Base, the gap is 4.17 points. Crucially, the gains extend to out-of-distribution benchmarks — the paper reports best-in-class results on all evaluated OOD tasks at both scales, which is the real test of whether you've reduced spurious dependence or just overfit harder. The method also works at inference time without any training: suppressing high-drift token candidates during decoding improves accuracy on its own. Architecturally, SCAPO lives in the policy-gradient family (specifically GRPO's clipped-ratio framework), but its novelty is entirely in the credit-assignment mechanism. It generates semifactual prompts — surface-level paraphrases that preserve the mathematical content — computes token-level KL divergence between original and semifactual response distributions, converts these into normalized stability scores, and uses them to scale down per-token advantages for unstable tokens during early training. Stability alone gets no bonus; only instability is penalized. The compute overhead is the cost of generating semifactual prompts and running forward passes under them, which the paper doesn't quantify explicitly. The integrity picture is decent but not airtight. The benchmarks are community standards (AIME, AMC, MATH-500, Minerva Math, OlympiadBench), which is strong. Baselines include GRPO, DAPO, and Dr.GRPO — current methods, not stale ones. Code is released on GitHub. But there's no pre-registration, no independent replication, and the semifactual generation process itself could introduce confounds: how do you guarantee that a paraphrase truly preserves the underlying problem? The paper acknowledges this implicitly by restricting to math, where semantic equivalence is more verifiable, but the boundary is fuzzy. The successor question is obvious: does this work on non-math reasoning, and does it scale to larger models? Math is the easiest domain for semifactual interventions because you can verify semantic equivalence relatively cleanly. Code, logic, science QA — these are harder. The authors stayed at 1.7B and 4B parameters. My read: they're saving the 7B+ experiments for the next paper, and the non-math domains are genuinely harder to crack because semifactual generation becomes less reliable. The compute cost of generating and evaluating semifactuals at scale is also non-trivial, and the paper is quiet about wall-clock overhead — likely because the numbers aren't flattering. What makes this paper worth your time isn't the magnitude of the gains — 4–6 points on AIME is meaningful but not paradigm-shifting. It's the diagnostic framework. The semifactual intervention lens gives you a way to audit any RLVR-trained model for spurious prompt dependence, and the token-level stability scores are a reusable tool. If you're building reasoning systems with GRPO or its variants, this is a concrete, implementable improvement with released code. If you're watching the RLVR space more broadly, the real question is whether causal credit assignment becomes standard practice or stays a niche technique.