Imagine you're learning to cook and you burn the garlic. You don't need a chef to physically guide your hand — you already know how to adjust the heat, how to scrape a pan, how to restart a sauté. What you need is someone to glance at the situation and say "turn the heat down and start over with fresh garlic." The correction isn't a new skill; it's a new composition of skills you already have, triggered by someone who can read the scene. That's the core mechanism of skill-space shooting. The committed claim: robots can autonomously improve their own task policies by using foundation models (vision-language models) to search through a library of reusable short behaviors — skills — to find corrections for failures, then distill those successful corrections back into the policy. No human demonstration of the correction is needed. The foundation model reads the scene, proposes which skills to try, and the robot runs the trials. Successes become training data for the policy itself. The architecture sits at the intersection of hierarchical reinforcement learning, model-predictive control (the "shooting" in the name comes from shooting methods in trajectory optimization), and foundation-model-as-planner agentic systems. The key structural choice is to search in skill space rather than action space — instead of optimizing over raw motor commands, you optimize over sequences of named, reusable skills. This dramatically reduces the search dimensionality and makes the search legible to a language model, which can reason about "grasp" and "push" but not about joint torques. The foundation model serves as both a proposal distribution (which skills to try) and a scene evaluator (did the correction work?). The ladder here is tricky. The paper positions itself against the standard paradigm of learning from human demonstrations (LfD) and recent agentic task-completion systems. The honest comparison is: most robot policy improvement methods require either (a) more human demos for each failure mode, (b) reward engineering, or (c) expensive online RL. Skill-space shooting sidesteps all three by using the foundation model as a free-but-noisy supervisor. The paper reports repeated real-world improvement on autonomous tasks, and — critically — shows that skills transfer across tasks, reducing the teaching burden for new tasks. But we're missing hard numeric baselines: no direct comparison to, say, DAgger, ReST, or other self-improvement baselines with matched compute budgets. Integrity is the main concern. The validation is real-world experiments, which is a strong signal — physical experiments are harder to p-hack than simulated ones. But the experiments appear to be conducted entirely by the authors' own lab, on their own tasks, with their own metrics. No community benchmark, no pre-registered task suite, no independent replication. The selection of tasks and the definition of "improvement" are controlled by the same team making the claim. The paper does show videos and promises a project page with additional results, which helps transparency. The milestone question is where this gets interesting. The current demonstration is repeated policy improvement on a handful of manipulation tasks with a finite skill library. The next concrete threshold is: can this work with 50+ skills across 20+ tasks in a shared household-scale environment, where the skill library is itself growing? That's roughly the scale at which you'd trust a home robot to self-improve without supervision. We're probably 3-5 years away, gated by foundation model reliability and physical robot deployment hours more than by the algorithmic idea itself. The obvious experiment the authors didn't run: scaling to a much larger skill library (100+ skills) where the foundation model's proposal distribution becomes the bottleneck, and measuring how proposal quality degrades. My read is (a) — they ran out of robot hours. Physical experiments are expensive and slow, and demonstrating the core mechanism on a few tasks is the rational move for a first paper. The next paper will almost certainly be the scaling study.