Imagine you're training a new sous-chef. You could teach them 200 specific recipes, or you could teach them the five mother sauces and how heat transforms protein. The second approach is smaller in scope but transfers everywhere — that chef can improvise when they see an unfamiliar ingredient. WOVEN makes the same bet for vision-language AI: instead of training models on each downstream task separately, teach them the underlying skill of predicting what happens when things change in visual scenes, and watch that skill propagate outward. The committed claim: visual transition reasoning — the ability to predict, explain, and reason about changes between visual states — is a single learnable primitive that transfers across spatial, embodied, physical, and temporal reasoning tasks. This is not another benchmark paper. It's an argument that the field has been treating symptoms (poor spatial reasoning, poor physics reasoning, poor temporal reasoning) when the disease is one thing: models cannot mentally simulate what happens next. To test this, the authors built a structured dataset of 36,076 examples spanning 20 scene types, 5 action categories, and 8 reasoning types, generated from video-pretrained diffusion models to ensure visual realism. They then evaluated 38 frontier MLLMs — including GPT-5.4 and Qwen3-VL-235B-A22B — and found the deficit is not a scaling problem. The strongest models still fall far below human performance, and the gap persists across model families and parameter counts. This is the most important finding: scale alone does not fix this. The training results are where things get genuinely interesting. Fine-tuning on subsets of roughly 2,000 WOVEN examples each improved performance on 22 of 26 external benchmarks, with gains up to 27.3 percentage points. Even more striking: WOVEN data can substitute for 30–50% of a task's own training data while maintaining comparable accuracy. This is the hallmark of a genuine shared primitive — it compresses. You can train less on the specific thing because the general capability carries the load. The training recipe that emerges from controlled ablations is precise and actionable: select your training supervision by the reasoning operation it teaches (comparison, prediction, counterfactual), not by the surface domain (kitchen scenes, driving, robotics). And prefer examples with larger visual state changes — subtle transitions teach less than dramatic ones. This recipe was validated prospectively on held-out benchmarks, not post-hoc on the same data, which matters enormously for credibility. Architecturally, WOVEN sits in the visual instruction-tuning family that traces from LLaVA through InternVL to the current crop of multimodal LLMs. The key structural choice is treating transition reasoning as a supervision signal layered on top of existing MLLM architectures at multiple scales, rather than proposing a new architecture. This is a smart constraint — it means the results are immediately applicable to whatever MLLM you're already using. The integrity picture is solid but not airtight. The 26 external benchmarks provide real breadth, and the prospective validation on held-out benchmarks is a genuine methodological strength. The main gap: the training data itself comes from video-pretrained generative models, not from real-world video or physics engines, which means there's an open question about whether the visual transitions are physically faithful or merely visually plausible. The authors acknowledge this implicitly by choosing diverse scene types, but the simulation-to-reality gap remains untested.