Imagine you're coaching a chess grandmaster. You can't touch their hands or see their internal calculations — you can only whisper advice before each move. After the game, you review the tape and think about what you should have said differently. But here's the catch: some of your "better" suggestions wouldn't have changed the grandmaster's play at all, and practicing those phantom corrections can actually make your future coaching worse. AdviSD is a method for figuring out which corrections to keep and which to throw away. The committed claim: a small (8B-parameter) language model can serve as a natural-language advisor to a frozen frontier LLM executor, and learning WHICH self-generated corrections to distill from — rather than distilling from all of them — yields meaningful gains over pure reinforcement learning. The paper proves formally that corrections whose targets don't strongly favor useful advice can limit learning in shared-parameter models, then builds a selection mechanism around that insight. The method pairs GRPO-style reinforcement learning with a selective self-distillation step. After an interaction, a feedback-conditioned copy of the advisor proposes corrections. The advisor then scores the same recorded executor response with and without its original advice — and the magnitude of that difference determines whether the correction gets used for training. This is clever: it measures whether the advice actually mattered to the executor's behavior without requiring executor logits or additional rollouts from the (expensive, frozen) frontier model. On the ladder, AdviSD with Qwen3-8B advisors beats advisor-GRPO by 4.2–6.4 percentage points on BFCL-v3 (function-calling benchmark) and 3.9–5.1 score points on EnvScaler (multi-step environment interaction). These are meaningful deltas for an 8B model advising much larger Gemini and Claude executors. The paper also shows the method beats random selection with matched correction counts, isolating the selection rule's contribution from the sheer effect of more training signal. The generalization results are the most interesting part. Trained advisors transfer across executor versions and model families — an advisor trained on one Claude variant works on a different one, and transfers to Gemini. This suggests the advisor is learning something about task structure, not just executor-specific quirks. Out-of-domain task transfer is also demonstrated, though the paper is lighter on quantifying how far that transfer stretches. Integrity is mixed. The benchmarks (BFCL-v3, EnvScaler) are community-recognized, and the ablation against random selection is the right experiment. But the evaluation is entirely the authors' own runs, there's no code release mentioned, and the formal proof — while clean — applies to a simplified setting. The selection mechanism's sensitivity to hyperparameters (the score-difference threshold) isn't deeply explored. Pre-registration is absent, as is standard for ML. The bigger picture: this sits in the fast-growing "LLM-steers-LLM" space where the central tension is whether you can get meaningful capability gains from orchestration without fine-tuning the executor. AdviSD says yes, and adds the useful nuance that selective learning from self-reflection outperforms blanket self-reflection. The next milestone is scaling this to harder multi-turn reasoning tasks and seeing if the advisor-executor gap narrows as executors get smarter — or whether the advisor becomes more valuable as a specialization layer.