Imagine you're playing Scrabble and you know the highest-scoring word on the board, but your tile bag keeps giving you mediocre draws. You could draw 64 times and pick the best set — that's power sampling — or you could rig the bag so the first draw is almost always great. OPPD rigs the bag. The paper's committed claim: a language model can be trained to produce, in one forward pass, answers that match or beat what you'd get from drawing 64 candidates and selecting the best under a power-sharpened distribution — and it does this without any reference answers or reward labels. The mechanism is a sequential Monte Carlo (SMC) sampler where the student model being trained generates candidate completions, a frozen copy of itself scores them under a power distribution (each completion's probability raised to an exponent β > 1, then renormalized), and those power-weighted probabilities become the loss weights for a maximum-likelihood update. The student trains on its own outputs, sharpened by its own frozen teacher. No external verifier. No reward model. The key mathematical insight is that raising completion probabilities to a power concentrates mass on the model's own most-likely answers — and that this sharpening can be distilled into the weights rather than paid for at inference time. The ladder is well-constructed. Against the untrained base model at the same temperature, OPPD adds up to +23.0 on MATH500 and +27.3 on GSM8K. More impressively, a single OPPD generation beats published power sampling with 64 candidates by 2.4 and 3.5 points on those benchmarks, recovering 94% of the gain that 16-candidate sampling gives the base model. Against GRPO — the current favored RL-from-verified-rewards approach — trained from the same checkpoint with the same compute budget and no reference answers, OPPD scores +3.8 on MATH500, +4.0 on GSM8K, and +5.4 on AIME. Crucially, the two methods are complementary: OPPD applied after GRPO adds up to 9.3 additional points. The integrity profile is solid but not airtight. MATH500, GSM8K, AIME, and HumanEval are community-standard benchmarks — no bespoke evaluation sets. Results hold across model families and sizes, including models already RL-trained, where simply lowering temperature gives nothing but OPPD still adds 4.4 points on MATH500. Code is released. The honest concern: all validation is self-reported, no independent replication exists yet, and the authors chose well-trodden math benchmarks where the model's own confidence is a reasonable proxy for correctness. On tasks where the model is confidently wrong — hallucination-prone domains — power sharpening could amplify errors. The paper doesn't test this. The sharpening exponent is controllable via a single loss coefficient, moving the effective β the model absorbs between 1.19 and 2.02 (versus 1.14 for ordinary on-policy distillation). This is a meaningful practical handle: it means you can tune how aggressively the model concentrates probability mass. The exponent rises mostly on the model's own high-scoring answers, confirming the method sharpens rather than reshapes the distribution. The transfer result is notable: trained only on mathematics, OPPD raises HumanEval (code generation) accuracy by up to 5.3 points. This suggests the method isn't just memorizing math patterns — it's learning a general sharpening behavior that transfers across reasoning domains. Whether this holds for domains further from math (e.g., open-ended writing, multi-turn dialogue) is the obvious untested question. The bigger picture: OPPD sits in the live argument about whether inference-time compute (sampling many candidates, tree search, chain-of-thought rollouts) should be baked into weights or spent at serving time. This paper is a strong data point for the 'bake it in' camp, showing that a substantial fraction of test-time search can be absorbed during training with no reward signal at all. The practical implication is cheaper deployment — one forward pass instead of 16-64 candidates — and composability with existing RL methods.