Imagine you're learning to cook a new dish by watching yourself try it on a mental instant replay. You picture yourself cracking eggs, flipping pans, plating food — all in your head. Most of your imagined attempts are garbage: you drop the pan, burn the eggs, plate a mess. But a few of those mental rehearsals actually look right, and you can verify them by asking: does the finished plate look correct? And do the hand motions I imagined actually match what happens in the video I pictured? The ones that pass both checks become your new training data. That's EVO-WAM in a single paragraph. The committed claim: a world action model (WAM) — a model that jointly generates future video frames AND the robot actions to produce them — can be adapted to entirely unseen manipulation tasks through iterative self-training on its own generated rollouts, without ever executing a single candidate action in a real or simulated environment. The paper calls this "evolving" the WAM, and the core trick is a two-stage verification pipeline: a vision-language model (VLM) checks whether the generated video actually depicts task completion, and an inverse dynamics model (IDM) checks whether the predicted actions are consistent with the video frames. Only trajectories passing both filters become training data for the next iteration. The architecture sits in the video-generation-as-policy family — a lineage running from SuSIE and UniPi through Cosmos and DreamZero. These are foundation-model-scale video generators repurposed as robot controllers. EVO-WAM doesn't replace these base WAMs; it wraps them in a self-improvement loop. Two specific augmentations make the loop work: state prediction heads (predicting robot proprioception alongside video) and anchored multi-frame context (conditioning generation on multiple past frames to prevent drift during long autoregressive rollouts). The compute regime is moderate — iterative fine-tuning of existing WAMs, not training from scratch. The ladder is where this gets interesting. On seven unseen RoboTwin 2.0 tasks in simulation, EVO-WAM lifts Cosmos3's average success rate from 26.9% to 68.0% (~2.5×) and DreamZero's from 28.5% to 46.4% (~1.6×). These are not small deltas. The real-world results are even more striking: on three long-horizon composite tasks, Cosmos3 goes from 20.0% to 76.7%, a gain of 56.7 percentage points. The baselines are the unmodified WAMs themselves, which is the right comparison — the question is whether self-training adds value, not whether this beats a completely different architecture. No comparison to offline RL or imitation learning baselines on the same tasks, though, which leaves a gap. Integrity is mixed. The simulation benchmark (RoboTwin 2.0) is a community benchmark, which is good. Real-world experiments on three tasks with 10 trials each (30 per method per iteration) are small-sample but meaningful as a proof of concept. The iterative training protocol means the authors saw intermediate results before deciding how many iterations to run and which tasks to report — a mild cherry-picking risk. No pre-registration, no code release mentioned at submission time (though a project page exists). The VLM and IDM verifiers are not independently validated against human judgment, so there's a bootstrapping concern: the quality of self-training is bounded by the quality of the verifiers. The milestone lens here is about closing the gap between generated-video supervision and human-demonstration supervision. At 68% success on unseen sim tasks, you're in the useful-but-not-deployable range. The target is 90%+ on diverse unseen tasks without any human demos, which would make WAM self-improvement a genuine alternative to data collection. The gap is probably 1-2 more architectural iterations plus better verification models. The obvious experiment not run: scaling to significantly more diverse unseen tasks (beyond 7 sim + 3 real), and stress-testing how the VLM verifier degrades as tasks get more semantically ambiguous. The honest read is (a) compute and time constraints — iterative training is expensive — combined with (c) saving broader task diversity for the next paper.