Imagine you're training a new driving instructor. The standard approach: hand them a checklist — 'say these things, ask these questions, use the Socratic method.' The Sherpa approach: put them in a car with five very different student drivers — one who panics at roundabouts, one who already knows theory but freezes in traffic, one who learns by doing — and score them purely on whether each student passes the test. The checklist never enters the picture. The instructor figures out what works by watching who learns. That's the core mechanism here. Sherpa is a multi-turn reinforcement learning framework that trains an LLM teacher not against predefined pedagogical criteria ('good teaching looks like X') but against the actual downstream learning outcomes of simulated students. The students are themselves LLMs, conditioned on distinct 'archetypes' — different learning preferences, knowledge gaps, and interaction styles. The teacher's reward signal comes entirely from whether these student archetypes demonstrate improved problem-solving after the tutoring session. This is a meaningful inversion: most prior work on LLM-as-tutor uses demonstrations of good teaching or preference data annotated by humans who know what teaching should look like, not what it accomplishes. The results are substantial on the metrics reported. Students instructed by Sherpa-trained teachers improve by an average of 20.5 percentage points across all archetypes. On MathTutorBench, the pedagogy score jumps from 52.5% to 79.2%. In human studies, the trained teacher is preferred over the base model 79.6% of the time in pairwise comparisons. These numbers span simulated and human evaluation, which is more than many papers in this space deliver. The architectural choice matters: this is multi-turn RL, not single-turn RLHF or supervised fine-tuning on teacher transcripts. The teacher model must plan across multiple dialogue turns, adapting mid-conversation to signals from the student. The student archetypes are the key design decision — they turn what would be a single-objective RL problem into a multi-objective one where the teacher cannot simply learn one strategy and apply it uniformly. The paper reports six distinct archetypes, though the abstract doesn't enumerate all of them. The conditioning mechanism (LLMs prompted to simulate different learner types) is pragmatic engineering, not a deep cognitive model, but it creates enough diversity to force adaptive behavior. The integrity picture is mixed. MathTutorBench is a community benchmark, which is good — it prevents the authors from grading their own homework on the pedagogy dimension. The human study (79.6% preference) adds a crucial external signal. But the student archetypes are themselves LLMs, which means the teacher is being optimized to teach machines to answer math questions, not to teach humans to understand mathematics. The transfer from simulated-student-passes-test to real-human-learns is the load-bearing assumption, and it's tested only indirectly through preference ratings, not through actual human learning outcomes measured pre/post. The ladder position is honest: the paper names its baseline (the base model before Sherpa training) and shows clear improvement. The MathTutorBench comparison gives a community-standardized reference point. What's missing is comparison against the strongest existing tutoring systems — Khan Academy's AI tutor, Khanmigo, or other RL-trained pedagogical agents. The 52.5% → 79.2% jump on MathTutorBench is striking, but without knowing where other systems land on the same benchmark, the ladder is incomplete. The gap that matters most: Sherpa has not been tested in a real classroom with real students and real learning outcomes measured over time. The human study measures preference (which teacher response do you like better?), not learning (did you actually understand the concept afterward?). This is the obvious next experiment, and the honest read is that it wasn't run because human learning studies are expensive, slow, and require IRB approval — classic resource constraint, not a red flag. But it means the paper's strongest claim ('paving the road towards AI tutors teaching real students') remains aspirational until that experiment happens.