Imagine you're teaching someone to cook by handing them recipes. The standard approach: give them a celebrity chef's recipes (off-policy expert data) and hope they learn the principles, not just memorize the steps. The problem is, those recipes assume skills your student doesn't have — so they copy the surface and miss the logic. Reinforcement learning takes the opposite tack: let the student cook freely, taste the results, and learn from their own successes. Better generalization, but the student has to stumble into something edible first, which can take a long time and a lot of wasted ingredients. This paper asks: what if you took the chef's recipes and progressively adapted them to your student's current skill level — not dumbing them down, but translating them into sequences of steps the student could plausibly have produced themselves? The core mechanism is a Markov chain Monte Carlo (MCMC) sampling algorithm that takes off-policy expert demonstrations and iteratively transforms them toward the on-policy distribution of the model being finetuned. Instead of modifying the loss function (the standard RL move), the authors modify the data distribution. The reference model generates candidate edits to expert traces, and the MCMC chain accepts or rejects these edits based on how likely the reference model would be to produce the resulting trajectory. Over iterations, the training data drifts from 'what the expert would do' to 'what a slightly better version of the current model would do' — bridging the distribution gap that makes vanilla SFT brittle. The claim, stripped bare: MCMC-reshaped data lets SFT match or beat on-policy RL methods (including strong baselines like ReST, RLHF variants, and rejection sampling) on scientific skill acquisition, mathematical reasoning, and open-ended expertise tasks — while forgetting less and generalizing better. This is a direct challenge to the conventional wisdom that RL is necessary for strong posttraining generalization. The paper does not argue that RL is useless, but that the SFT vs. RL gap is substantially a data-distribution problem, not a learning-objective problem. Architecturally, this sits in the family of data-centric posttraining methods — closer to rejection sampling and self-play than to policy-gradient RL. The MCMC chain is the key structural innovation: rather than filtering expert data (rejection sampling) or generating entirely new data (self-play), it continuously deforms existing expert traces toward on-policy plausibility. The compute cost is in sampling: each MCMC step requires model inference to score candidate trace edits. This is cheaper than full RL rollouts but more expensive than vanilla SFT on static data. The method is gradient-free at the data-shaping stage, with standard cross-entropy SFT applied to the reshaped data. Integrity-wise, the authors compare against strong on-policy baselines across multiple task domains — not just one cherry-picked benchmark. The claims about reduced catastrophic forgetting and better distributional performance (not just mode-seeking) are tested, and the paper acknowledges that the MCMC chain adds compute overhead. What's missing: no pre-registration, no independent replication, and the benchmarks — while spanning multiple domains — are author-selected. The strongest possible test (head-to-head against frontier RL pipelines at scale on established community benchmarks like MATH or MMLU) is not fully reported, leaving room for the result to shrink at production scale. The field fight here is real: posttraining research has been polarized between 'SFT is cheap but fragile' and 'RL is expensive but generalizes.' This paper argues the dichotomy is false — that sampling is the missing primitive that lets SFT absorb the strengths of RL without the instability. If the result holds at scale, the practical implication is large: SFT pipelines are simpler to build, debug, and parallelize than RL pipelines. The next milestone is demonstrating this approach at frontier model scale (70B+ parameters) on community-standard benchmarks with independent replication. The authors likely did not run that experiment because of compute constraints — not because it failed.