Imagine you're a locksmith who sells high-security padlocks, but your threat model only considers thieves who try to pick the lock — not thieves who make a wax impression of the key and then file it smooth at home. That's the core problem this paper identifies in current distillation defenses for large language models. Defenders test whether their countermeasures survive distillation, then declare victory. They never check what happens when the attacker spends another afternoon with reinforcement learning. The committed claim: distillation defenses that appear effective under the standard evaluation regime — test immediately after distillation — break when attackers apply even basic reinforcement learning (RL) as a follow-up step. The paper argues the entire threat model used by existing defenses is misspecified. You're measuring lock-picking resistance when the real attack is key-copying followed by filing. The mechanism is straightforward and that's what makes it damaging. Attackers collect reasoning traces from closed-source frontier models via standard APIs — not jailbreaks, not exploits, just normal usage. They distill a smaller model on these traces. Current defenses degrade these traces (add noise, truncate chain-of-thought, etc.) so the distilled model performs poorly. But the paper shows that RL on verifiable-answer tasks recovers the reasoning capability, effectively smoothing out whatever corruption the defense introduced. Simple attacks using only final-answer API outputs, combined with post-distillation RL, match sophisticated attacks that extract full hidden reasoning traces. The ladder placement is revealing. The paper doesn't compete against a classical baseline in the traditional sense — it's an attack paper, so the baseline is existing defenses. It names specific defense strategies (trace corruption, output perturbation) and demonstrates they fail under the extended threat model. The key finding is equivalence: cheap attacks plus RL reach the same performance as expensive full-trace extraction attacks. This collapses the cost curve for attackers. Integrity-wise, the experimental design tests the right thing: does RL undo defense degradation? The authors use existing closed-source models and publicly accessible APIs, which grounds the attack in realistic conditions rather than laboratory setups. However, the paper lacks independent replication and doesn't specify pre-registration. The choice to demonstrate on reasoning benchmarks is natural but also favorable terrain — RL with verifiable answers is the easiest RL setting. Whether the result extends to fuzzier capabilities (creative writing, nuanced judgment) is unstated. The practical implication is stark: any distillation defense that leaks enough signal to reconstruct approximate reasoning traces is dead on arrival. The paper's conclusion points toward batch-level defenses — analyzing patterns across many queries rather than corrupting individual outputs — as the more promising direction. This reframes the defense problem from signal degradation to behavioral detection, a much harder but potentially more robust approach. The successor question looms large. The authors did not test batch-level defenses themselves, nor did they demonstrate the attack at frontier scale (e.g., distilling from GPT-4o or Claude 3.5 into a competitive open model with full benchmark comparisons). The honest read: batch-level defenses are a different paper requiring different infrastructure, and frontier-scale demonstrations would require significant compute and potentially create dual-use concerns. But until someone runs those experiments, we're left with a strong negative result — current defenses don't work — without a proven replacement.