Imagine you're a chef with a massive spice rack — 50 jars, most unopened. A recipe calls for "season to taste," so you dump in everything. The dish is worse than if you'd used five spices chosen carefully. That's the core mechanism here: most inference-time "skills" injected into LLMs during distillation are noise or actively harmful, and the paper's contribution is a principled way to identify the useful minority. The committed claim: fewer than 25% of semantically retrieved skills produce useful distillation signals, and a method called SGUID (Selecting a compact skill bank for model-skill co-evolution) can pick a tiny subset — as few as 6 skills — that matches or exceeds distillation from the entire bank across four models from the OLMo and Qwen families. This isn't a retrieval improvement. It's a filtering mechanism that asks whether each skill actually teaches the student model something during on-policy rollouts. SGUID works by retaining a skill only if it consistently yields effective learning signals during training — essentially measuring whether the teacher's skill-conditioned policy actually moves the student in a useful direction. The selected skills are distilled, the model is updated, and then a new candidate bank is curated from the updated model's rollouts. The loop repeats. In round two on Qwen3-8B, 3 newly selected skills pushed performance from 64.3% to 66.3%. The full banks being compared against are up to 11× larger. The stability result is arguably more important than the efficiency result. On Qwen3-4B, naively updating the model with unfiltered skills degrades performance — including a 0.3 percentage-point drop on HMMT25. SGUID improves HMMT25 by 0.5 points after round one and 1.1 points after round two. Without selection, the co-evolution loop is unstable; with it, it converges upward. This is the paper's strongest empirical contribution. The evaluation uses mean avg@12 across benchmarks on four models (OLMo and Qwen families, including Qwen3-4B and Qwen3-8B). The baseline is full-bank distillation — the standard approach from prior work (Li et al., 2026). On three of four models, 6 selected skills match or beat the full bank in round one. After round two with 3 additional skills, all four models match or exceed full-bank performance. The HMMT25 benchmark provides the most granular stability signal. The paper sits in the growing "skills as inference-time patches" literature, where reusable procedural guidance is retrieved and injected at test time. The field fight is about whether scaling the skill bank is the right move or whether curation matters more. SGUID lands firmly on the curation side — more skills aren't just wasteful, they're destabilizing. The analogy to curriculum learning in education is direct: not every lesson is worth teaching at every stage of a student's development. What's missing is any test beyond two rounds of co-evolution and any evaluation outside the OLMo/Qwen families. The paper also doesn't address what happens when the skill bank itself is generated adversarially or from weaker models. The iterative loop is promising but undertested — we see a proof of concept, not a scaling law.