Imagine you're a flight instructor who needs to test whether a student pilot can handle every possible landing scenario. You can't fly every scenario for real — too dangerous, too expensive. So you build a simulator. The question is: how faithful does that simulator need to be, and how do you prove your student won't crash in a scenario the simulator never showed? That's the core problem in verifying vision-based neural controllers. A self-driving car or emergency braking system uses a camera feed processed by a neural network. To formally verify that this system won't fail, you need a model of what the camera could possibly see — a perception surrogate. GANs have been the go-to surrogate, but they're enormous, reproduce complex scenes poorly, and resist the kind of formal analysis verifiers need. This paper proposes swapping GANs for stochastic world models with physically grounded latent variables — and then develops a verification procedure that actually works on the resulting system. The committed claim: a world model built from verifier-friendly operations (operations whose output bounds standard interval/symbolic verifiers can compute) reproduces held-out camera frames more faithfully than GAN surrogates that have up to 130× more parameters. That's not a marginal improvement — it's a different regime of model efficiency. The world model's latent space is anchored to physical quantities (position, velocity, lighting conditions), which means the verifier can reason about how changes in the physical state map to changes in pixel observations. The verification procedure is the second load-bearing contribution. It combines four techniques: falsification (find counterexamples fast), adaptive refinement (zoom into ambiguous regions), symbolic analysis (propagate bounds through the network), and backward analysis (trace from unsafe outputs back to inputs). On an existing emergency braking benchmark using a GAN surrogate, this procedure resolves 100% of the state space — the prior state-of-the-art verifier left 38% unresolved. That's not incremental; the unresolved fraction went from 38% to zero. The RGB version of the benchmark is where the paper enters genuinely new territory. No verification results had previously been reported for the full-color version. The world model surrogate, combined with the new verification procedure, resolves over 80% of the RGB state space. This is a first result, not a SOTA comparison — there's no prior number to beat because nobody had a number at all. The architectural bet is clear: physically grounded latents over learned-but-opaque GAN representations. This trades expressiveness for tractability. A GAN can hallucinate arbitrary scenes; a physics-anchored world model is constrained to scenes that are physically plausible. For verification, that constraint is a feature — it shrinks the space the verifier must search. But it also means the approach is tethered to domains where you can write down the physics. Emergency braking on a straight road is one thing; verifying a controller navigating a crowded urban intersection is a different beast entirely. The integrity story is mixed. The benchmarks are from existing literature (emergency braking scenarios used in prior verification work), which is good — they didn't invent a benchmark they'd win. But all validation is computational, with no physical experiment. The 130× parameter reduction is striking, but the baseline GAN surrogate's fidelity isn't independently validated. The 38% resolution improvement is measured against a named prior verifier, which grounds the comparison. Code and reproducibility status are not stated in the abstract. The milestone that matters is scaling: can this approach handle perception surrogates for scenes with more than a handful of physical variables? Emergency braking on a track is maybe 5-10 latent physical dimensions. Urban driving might require hundreds. The 80% RGB resolution is the number to track — if subsequent work pushes that toward 95%+ on more complex scenes, this approach becomes a real pipeline component for autonomous vehicle certification.