Imagine you're driving in a city you half-know. You can glance in the rearview mirror all day, but that mirror only helps if you've internalized how traffic flows — which lane merges, which light sequence repeats. A driver who learned routes by watching dashcam footage forward in time develops a causal intuition: this scene leads to that scene. A driver who only studied still panoramas has no such instinct. The mirror is useless to them. That's the core mechanism here: giving a robot more visual history is worthless unless the video foundation model was trained to predict the future from the past — autoregressively, frame by frame, cause then effect. Long-WAM's committed claim is that autoregressive (AR) video pretraining is a prerequisite for robots to benefit from longer observation context under real-time control. The paper backs this with a clean ablation: on RoboCasa GR-1 tasks, extending context from 0 to 19.2 seconds with an AR-pretrained backbone pushes success from 63.3% to 78.7%. The same extension with a bidirectional (masked) pretrained backbone shows no net gain. That's the headline finding — not just that longer context helps, but that it only helps with the right inductive bias baked into pretraining. The architecture is a causal world-action model built on top of a video generation backbone. Phase one pretrains on robot and egocentric video without action labels, learning next-frame prediction autoregressively. Phase two adapts this to world-action modeling, preserving the causal history-to-future structure while conditioning on actions. The design leans on streaming observation encoding — processing frames as they arrive rather than batching — plus asynchronous action execution that decouples inference latency from control frequency. On an RTX 5090, each action chunk including future-video latent prediction takes 107.4 ms, which is tight but viable for real-time humanoid control. The ladder positioning is strong. On LIBERO-Long, RoboTwin 2.0, and DOMINO, Long-WAM achieves best results among compared methods. The most striking comparison is against Pi0.5 and Fast-WAM on dynamic cup stacking with a Unitree G1 humanoid: Long-WAM hits 95% success across 20 trials where both competitors score 0%. Robot-domain AR pretraining further improves peak success on GR-1 and LIBERO-Long beyond generic AR pretraining, suggesting domain-matched causal video data compounds the benefit. Integrity is mixed. The benchmarks are community-recognized (RoboCasa, LIBERO-Long, RoboTwin 2.0, DOMINO), and the ablation cleanly isolates the AR vs. bidirectional pretraining variable — that's the strongest card. But evaluation is all same-team, the 20-trial sample on dynamic stacking is small, and the 0/20 failures for Pi0.5 and Fast-WAM cry out for explanation of whether those baselines were configured at their best. Hardware deployment across three NVIDIA platforms (RTX 5090, DGX Spark, Jetson AGX Thor) adds credibility that the system isn't a benchmark-only artifact. The milestone question is about scaling context further and hitting harder manipulation tasks. Today we're at 19.2 seconds of usable visual history on humanoid platforms. The next unlock is probably 60+ seconds of context for genuinely long-horizon household tasks — think multi-room cleanup or cooking sequences that span minutes, not seconds. That likely requires either more efficient video tokenization or dedicated memory-compression layers, and it's probably 2-3 years out given current trajectory. The obvious experiment not run: scaling context beyond 19.2 seconds to find where the curve plateaus or breaks. The paper stops at a point where performance is still climbing. My honest read is (a) compute and time budget — longer contexts demand proportionally more memory and inference time, and pushing to 60+ seconds on current hardware would likely break the real-time constraint, especially on Jetson. There's also no systematic noise or failure-mode analysis: when Long-WAM fails at 19.2s context, what went wrong? That would strengthen the mechanistic story considerably.