Imagine you're teaching someone to cook by watching a chef. You could transcribe the chef's running commentary — "I'm searing because the Maillard reaction needs 350°F" — or you could just write down: "Step 3: Sear the steak. Step 4: Rest it." This paper asks whether the stage labels ("you're in step 3") are more useful training signal than the full chain-of-thought reasoning, and finds that they are — up to a point. The committed claim: a 1.7-billion-parameter student model, trained offline on 404 teacher demonstrations annotated with short task-stage labels, achieves 72.4% success on unseen ALFWorld household tasks — matching action-only distillation and dramatically outperforming a reasoning-trained student (48.3%) that uses constrained action selection. The mechanism is called Task-Progress Distillation (TPD). Each demonstrated action is paired with a compact label describing the current stage (e.g., "acquiring object" or "processing object"). The student scores admissible stage–action pairs jointly, and a deterministic harness executes the chosen action. The interesting result is the interaction between stage labels and data volume. At 200 demonstrations, TPD's stage labels boost success from 48.0% to 67.7% over action-only supervision — a nearly 20-point lift. But at 404 demonstrations they tie, and at 808 both converge to 76.9%. The stage labels function as a structured prior that compensates for sparse data; with enough examples, the student infers the latent task structure on its own. This is a clean empirical finding about the information geometry of distillation. Architecturally, this is a supervised fine-tuning approach on a 1.7B language model (the paper doesn't name the base model architecture explicitly, but the parameter count and the ALFWorld setting place it squarely in the small-LM agent family). The student operates in a score-and-select loop: given the current observation, it scores each admissible (stage, action) pair and picks the highest. A deterministic harness handles execution — the student never calls an API or reasons step-by-step at inference time. This is offline distillation with structured labels, not online RL or chain-of-thought prompting. The shared-history analysis is the most methodologically interesting piece. The authors isolate decision points where multiple trajectories share the same history prefix, then compare which approach picks the right branch. TPD's advantage concentrates at subgoal transitions — specifically, moving from object acquisition to processing. This is where a naive action-copier gets confused because the surface actions change category, but the stage label stays coherent. It's a tidy mechanistic explanation for when structured supervision helps. Validation is limited to ALFWorld, a text-based household simulator with six task types. This is a standard benchmark in the LLM-agent literature, but it's a single environment with limited action spaces. No real-world deployment, no transfer to other benchmarks (WebShop, ScienceWorld, etc.), and no comparison to the strongest recent agent frameworks like ReAct or Reflexion running on similarly-sized models. The 48.3% reasoning baseline uses constrained action selection, which may not represent the best possible reasoning-distillation approach. The convergence at 808 demonstrations is both the paper's most honest result and its biggest limitation. It tells you that explicit task-stage labels are a data-efficiency trick, not a fundamental capability unlock. If you can generate enough teacher demonstrations — which is increasingly cheap with large models — the stage labels become redundant. The practical question is whether the demonstration budget is the binding constraint in your deployment, and for many real-world agent applications, it may not be.