Imagine you're teaching someone to drive. You're an expert — you can parallel park blindfolded, merge at speed, read traffic three moves ahead. But if you grab the wheel and execute a maneuver the learner physically can't reproduce — a snap-judgment lane change they don't have the spatial awareness for yet — your demonstration is worse than useless. It's confusing. The best driving instructors don't show off. They give the next instruction the student can actually execute from where they are right now. That's the core mechanism of this paper. The committed claim: in on-policy self-distillation for language models, a privileged teacher that simply maximizes task reward can actively degrade the student, because it solves problems via shortcuts the unprivileged student cannot access. JOLT fixes this by adding a KL-regularization term that anchors the teacher to the student's current distribution, ensuring dense token-level guidance stays actionable. The paper derives a necessary and sufficient condition — the teacher's distillation gradient must be a positive scalar multiple of the student's reward gradient — and shows that unconstrained privileged teachers violate this routinely. The architectural setup is clean: a single policy model plays both roles. As teacher, it receives privileged information (ground-truth answers, tool outputs) and trains with outcome reward plus a KL penalty toward the student's current distribution. As student, it trains via standard on-policy distillation from the teacher's token-level probabilities. No separate teacher model, no frozen checkpoints. The KL term is doing the heavy lifting — it's what prevents the teacher from drifting into territory the student can't follow. This sits in the self-play / self-improvement family alongside STaR, ReST, and SPIN, but the theoretical contribution is the alignment condition between teacher and student gradients. The ladder comparison is where the paper earns respect. JOLT is benchmarked against GRPO (a strong RL baseline), standard self-distillation, and ablations across four domains: MATH-500, LiveCodeBench, tool use (BFCL), and terminal use (SWE-bench Verified). The paper reports that JOLT matches or exceeds GRPO performance while requiring fewer training samples to reach equivalent accuracy, with the gap widening on harder tasks. On MATH-500, JOLT achieves competitive pass rates; on coding and tool-use benchmarks, it shows clearer wins. The honest caveat: gains over GRPO are modest on easier tasks and more pronounced in the long-horizon, sparse-reward regime where standard RL struggles most. Integrity is reasonable but not bulletproof. Benchmarks are community-standard (MATH-500, LiveCodeBench, SWE-bench Verified, BFCL) — not cherry-picked toy tasks. However, all results are from the same team, there's no independent replication, and the paper doesn't state whether benchmark selection was pre-registered. The theoretical result (Theorem 1 on the gradient alignment condition) is a genuine mathematical contribution, but the empirical validation is simulation-only on the authors' own infrastructure. The ablation study — removing the KL term, removing student rewards — is well-designed and isolates the contribution cleanly. The milestone to watch: JOLT demonstrates efficiency gains at current model scales, but the real question is whether this teacher-student alignment condition scales to frontier models (100B+ parameters) and whether the KL regularization coefficient needs careful tuning per domain or generalizes. If JOLT-style training can cut sample complexity by 2-3× at frontier scale, it changes the economics of RLHF and post-training broadly. The gap from here to there is probably 6-18 months of scaling experiments. The obvious experiment not run: applying JOLT to RLHF with human preference data rather than outcome rewards. The paper uses verifiable outcome rewards (math correctness, code execution) where you can score trajectories cleanly. Human preference is noisier, and the KL alignment condition might behave differently when the reward signal itself is uncertain. My read: this is (c) — they're saving it. The theoretical framework generalizes, and preference-based JOLT is the natural next paper.