Imagine you're a bank auditor checking a company's books. The company reports a net cash flow of zero — money in equals money out. Looks clean. But if you take the absolute value of every transaction before summing, you discover $50 million sloshing around in suspicious directions. The net was zero only because a massive deposit and a massive withdrawal happened to cancel. That's the core mechanism of this paper: when LLM reinforcement learning systems check whether a generated response is still "on-policy" enough to learn from, the standard method averages signed log-probability ratios across tokens — letting a token that became 10× more likely cancel out a token that became 10× less likely. CARM takes the absolute value first, exposing the true magnitude of drift. The committed claim: sequence-level off-policy masking in LLM RL has a cancellation bug, and fixing it with absolute-value log-ratios yields a theoretically grounded mask that improves both math reasoning and code generation by 2-3 percentage points over the standard geometric-mean mask. This is not a new RL algorithm — it's a diagnostic fix to an existing pipeline component that most practitioners have treated as settled. The paper sits in the post-training RL stack that has become standard since DeepSeek-R1 and similar systems demonstrated that GRPO-style reinforcement learning dramatically improves LLM reasoning. The specific problem — rollout responses becoming off-policy between generation and gradient update — is well-known, and the geometric-mean importance ratio is the default solution in production systems. CARM replaces one line of that pipeline: the averaging operation. The authors prove that accepted responses under CARM satisfy a joint bound on both the fraction of token ratios outside a prescribed band and the mean distance of outlier ratios beyond that band. This is a tighter guarantee than the geometric mean provides, because the geometric mean can accept responses where half the tokens have drifted wildly in opposite directions. On the ladder: the baselines are geometric-mean masking (the industry default), clip-higher masking, and no masking at all. CARM improves mean@16 averaged over AIME 2024/2025/2026 and BeyondAIME by up to 3.13 percentage points over geometric-mean masking. On code, average pass@1 across four benchmarks rises 2.88 points over the strongest evaluated baseline. These are meaningful but not paradigm-shifting deltas — the kind of improvement that matters in production systems where every point counts, not the kind that rewrites textbooks. Architecturally, this is a gradient-based policy optimization method in the GRPO family — group relative policy optimization with verifiable rewards, no critic model. The masking operates at sequence level: a binary accept/reject decision per response, applied before the policy gradient step. The key structural choice is replacing the signed average with an unsigned (absolute-value) average of per-token log importance ratios. Compute overhead is negligible — one extra absolute-value operation per token. The method is agnostic to the underlying RL algorithm and could slot into PPO, REINFORCE, or any importance-weighted policy gradient. Integrity is reasonable but has clear limits. The benchmarks — AIME variants and LiveCodeBench, MBPP+, HumanEval+ — are community-standard and public. The base model is Qwen2.5-32B, a strong open-weight foundation. The authors run ablations on the threshold parameter and show sensitivity curves. However, there's no pre-registration, no independent replication, and the 28-page paper doesn't release code at submission time. The training data (57K math problems, 40K code problems with test cases) is described but its exact composition could influence results. The comparison is against masking variants, not against entirely different off-policy correction strategies like full per-token IS clipping. The milestone question is practical: at what scale of RL training (model size, rollout parallelism, update frequency) does cancellation-induced masking error become the dominant failure mode? The paper demonstrates the problem at 32B parameters. The next meaningful number would be demonstrating that CARM's advantage holds — or grows — at 70B+ scale where the gap between rollout and training policy widens with more parallel generation. The authors don't test this, most likely because of compute budget constraints — training 70B models with RL is expensive, and this reads as a focused paper from a team optimizing their 32B pipeline.