Imagine you learned to cook by watching 10,000 hours of cooking shows — you understand how eggs crack, how dough stretches, how oil splatters. You've never touched a specific kitchen, but you know what happens when things interact. Now someone hands you a new set of utensils and says 'predict what happens next.' You'd be surprisingly good — because the physics of cooking doesn't change when you switch from a Viking range to a hotplate. What changes is how your hands map to the tools. That's WorldLine's core trick: decouple the universal dynamics of manipulation (things slide, topple, deform) from the embodiment-specific control signal (how a particular robot arm moves through space). The committed claim: WorldLine is the first visual simulator that trains manipulation dynamics on action-free video at scale (10,000+ hours), grounds them across heterogeneous robot embodiments (10+ types, 2,000+ hours of action trajectories), and does so through a shared image-space action representation rather than requiring a common control space. This isn't 'we fine-tuned a video model on robot data.' It's a two-stage architecture where Stage 1 learns physics from passive observation and Stage 2 learns the mapping from actions to those physics — and the stages are separable. The architectural insight that makes this work is the image-space action representation. Instead of conditioning on joint angles or end-effector poses (which differ between a Franka and a UR5 and an AgiBot), WorldLine projects actions into pixel-space overlays that any video model can consume. This is the Rosetta Stone move: translate the tower-of-Babel problem of incompatible control interfaces into a lingua franca the dynamics model already speaks. Multi-view training and deliberate failure-enriched data (showing what goes wrong, not just what goes right) sharpen the model's sensitivity to interaction dynamics — the moments where robot touches object and something changes. On the ladder: WorldLine improves robot-mask IoU by 0.1626 over the strongest baseline on failed trajectories, which is the hardest test — predicting what a failing robot looks like is much harder than predicting success because failure is diverse and underrepresented in training data. It predicts trajectory success at 74% mean accuracy across RoboTwin and AgiBot benchmarks, edging the strongest baseline by one percentage point. The real headline number is the 21.4 percentage point improvement in task success over direct policy execution on RoboTwin — achieved with zero RoboTwin-specific training or adaptation. That's the out-of-domain generalization result that makes this more than a benchmark paper. The integrity picture is mixed but honest. The benchmarks (RoboTwin, AgiBot) are community-facing and multi-embodiment, which is good. The failure-trajectory evaluation is a deliberate stress test the authors chose to highlight, which suggests confidence rather than cherry-picking. But the validation is entirely simulation-to-simulation — WorldLine predicts video rollouts, and those rollouts are evaluated against ground truth video, not against real physical outcomes. The real-world deployment gap remains unaddressed. No pre-registration, and we don't know if baseline selection was pre-committed. The milestone to watch is straightforward: WorldLine currently predicts trajectories for evaluation and planning. The next unlock is closed-loop real-world deployment, where the simulator's predictions drive actual robot behavior in physical environments at interactive rates. The few-step distillation already targets inference speed, but the paper doesn't demonstrate real-time closed-loop control on physical hardware. That's the gap between 'useful planning tool' and 'replaces a physics simulator.' Given the 21.4pp improvement in sim, the question is whether that margin survives contact with real sensor noise, lighting variation, and actuation error. The experiment the authors did not run — and the one everyone will ask about — is real-robot closed-loop evaluation. They have the distillation pipeline for speed, they have the multi-embodiment grounding, and the planning results are strong. The most likely read: they're saving it for the next paper, not because it failed, but because real-robot experiments require hardware access and lab time that's orthogonal to the modeling contribution. The project page exists, the architecture is complete, and the sim results are the proof-of-concept needed to justify that next step.