Imagine you're driving a rental car for the first time. You don't know how sensitive the brakes are or how much the steering pulls left, so for the first few blocks you run little experiments — tap the brakes, nudge the wheel — and within minutes your brain has selected a mental model of this particular car from a library of cars you've driven before. You didn't retrain your entire nervous system; you just picked the best match from a shelf of prior models and started planning with it. That is exactly what LLA-MPPI does for a legged robot. The committed claim: a legged robot can adapt in real time to unknown physical changes — added payload, a disabled leg, altered contact friction — by selecting among a bank of GPU-batched contact simulators rather than learning a new model or relying on offline training. The method achieves 97.5% task success across four simulated scenarios, compared to 74% for the strongest adaptive baseline (L1-MPPI) and 98.5% for an oracle that knows the true model. The gap between LLA-MPPI and the oracle is 1 percentage point. The gap between LLA-MPPI and the next best method is 23.5 points. Architecturally, this belongs to the sampling-based MPC family — specifically Model Predictive Path Integral (MPPI) control, which replaces gradient computation with massively parallel trajectory rollouts scored by a cost function. The key structural move is the "look-back" step: a sliding window of recent state observations is replayed through every candidate simulator in the bank, and the simulator with the lowest prediction error gets promoted to "current model" for the "look-ahead" planning phase. The entire pipeline — bank evaluation, model selection, trajectory optimization — runs on GPU in parallel. No neural networks are trained. No offline data collection is required. The hypotheses are interpretable: you can inspect which simulator was selected and see that the robot has effectively diagnosed itself as "carrying 3 kg on my back" or "missing my rear-left leg." The ladder comparison is honest but narrow. The authors benchmark against vanilla MPPI (no adaptation), L1-MPPI (an L1 adaptive controller layered on MPPI), and an oracle MPPI that uses the true simulation parameters. Classical adaptive control methods like L1 require a model structure that rigid-body contact dynamics don't naturally provide — the authors argue this is a fundamental mismatch, not just a tuning failure. The 97.5% vs 74% gap on the four simulated tasks is convincing, but the task suite itself is limited: payload addition, leg disabling, box pushing with mass changes, and terrain friction changes. These are the canonical hard cases, but they're also the cases most favorable to a discrete hypothesis bank. Integrity is mixed. Hardware validation on a real Unitree Go2 — walking under a payload added mid-run, walking with a disabled leg, pushing a box with increasing mass — is a strong signal that the method isn't just a simulation artifact. But the four-task simulation benchmark was clearly designed around the method's strengths (discrete parameter changes that map neatly to a hypothesis bank). There's no pre-registration, and the baselines, while reasonable, don't include recent learning-based adaptive methods like RMA or rapid motor adaptation via domain randomization. Code and videos are released, which is good practice. The milestone question is about scale. The current bank size appears to be on the order of tens of hypotheses. The method's compute cost scales linearly with bank size (each hypothesis is a parallel simulation), so the real question is: can you scale to hundreds or thousands of hypotheses covering continuous parameter spaces without blowing the real-time budget? The authors discretize what is fundamentally a continuous parameter space. If the bank can grow to ~100-500 hypotheses with interpolation or hierarchical selection while maintaining sub-10ms planning cycles, you get a general-purpose adaptive whole-body controller that handles combinations of failures, not just single discrete changes. The obvious experiment not run: simultaneous multi-parameter changes where the correct model isn't in the bank at all. What happens when the robot loses a leg AND gains a payload AND hits unexpected terrain — a combination not represented by any single hypothesis? The authors almost certainly know this is the weakness. My read is (c): they're saving it for the next paper, likely with some form of hypothesis interpolation or online bank expansion. The current results are strong enough to publish without solving composition, and the compositional case is the natural sequel.