Imagine you're a chess player who can name every piece on the board but freezes when asked 'what happens after Nf3?' That's the gap SpaceCast-Bench exposes in today's vision-language models. They can describe what's in front of them — relative positions, containment, occlusion — but ask them to predict what a scene looks like after an object is moved, rotated, or removed, and performance craters. The benchmark draws a hard line between spatial perception (reading off what you see) and predictive spatial reasoning (simulating what you can't see), and the results are damning for anything claiming spatial intelligence. The committed claim: no existing benchmark directly and diagnostically evaluates predictive spatial reasoning in VLMs. SpaceCast-Bench fills that hole with 3,862 questions across 182 real-world scenes, organized into 16 task types at three difficulty levels — static perception, local prediction, and global prediction. The progression is deliberate: first confirm the model can read the scene, then ask it to update spatial state after an intervention, then require relational inference over the unseen outcome. It's an observe-transform-infer pipeline, and the 'transform' and 'infer' stages are where models collapse. Twenty-one models were evaluated. The best performer — not named in the abstract but presumably a frontier multimodal model — hits 58.0% accuracy. Humans score 87.2%. That's a 29.2-percentage-point gap on tasks humans find relatively straightforward. More striking: models specifically marketed as spatially specialized perform near random chance, which is roughly 25% on four-choice questions. This isn't a matter of scale or training data volume — it's a fundamental capability gap in how these models represent and manipulate spatial state internally. The controlled analyses yield two actionable findings. First, bridge views — intermediate camera angles that connect distributed observations of the same scene — are critical for models to integrate information across viewpoints. Without them, performance degrades sharply, suggesting models aren't building coherent 3D scene representations but rather stitching together 2D pattern matches. Second, when given explicit 3D evidence (point clouds, depth maps), models improve more reliably than when given generated outcome images or videos. The implication: hallucinated visual predictions from generative models introduce more noise than signal for downstream spatial reasoning. The fine-tuning result is the paper's quiet knockout punch. Taking Qwen3-VL-4B — a 4-billion-parameter model — and training it on programmatically generated SpaceCast data lifts accuracy from 34.0% to 65.7%, a 31.7-point jump that actually exceeds the best zero-shot frontier model. More importantly, this transfers: macro-average gains appear across six out-of-domain spatial benchmarks. The training data is synthetic and programmatic, meaning it can scale without human annotation, and the gains generalize, meaning the model is learning something about spatial reasoning rather than just memorizing task patterns. The paper sits in a growing argument about whether scaling language-model pretraining will eventually produce spatial intelligence or whether explicit 3D inductive biases are required. SpaceCast-Bench provides ammunition for the 'explicit structure needed' camp: raw scale hasn't closed the gap, specialized models haven't either, but targeted synthetic training with 3D-grounded data has. The benchmark is open, the code is released, and the dataset is on HuggingFace — all the right moves for a community resource. What's missing is a deeper ablation on what makes the synthetic training data work. Is it the 3D grounding? The diversity of transformations? The programmatic generation ensuring clean labels? The paper also doesn't evaluate the fine-tuned model on harder compositional tasks (multiple sequential transformations, adversarial distractors) or test whether the gains hold at larger model scales. These are natural next experiments, and the authors likely know exactly what they'd show — which is why SpaceCast-Bench v2 probably already exists in draft form.