Imagine you're training a new air traffic controller. During training, you let them sit next to a veteran who has a radar overlay showing exactly which blips matter for each decision. The trainee watches the veteran work with the overlay, then has to make the same calls without it. Over time, the trainee internalizes the veteran's selective attention — learning WHERE to look, not just WHAT to look for. That's the core mechanism of Where-OPD: give the teacher version of a multimodal LLM spatial coordinates telling it which image regions matter, let it reason with that privileged information, then train the student version to reproduce the same answers from the raw image alone. The committed claim: spatially grounded textual hints — not image crops, not external teacher models, not human annotations — can serve as the privileged information channel in on-policy self-distillation for multimodal LLMs, and the resulting perceptual improvements transfer from synthetic training scenes to real-world benchmarks. The prior art here (Zoom-OPD and similar approaches) relied on visual zooming — cropping relevant image regions and feeding them to the teacher. That works for tasks where zooming helps, but it's narrow. Where-OPD instead gives the teacher text telling it "the red cup is at coordinates (0.3, 0.7)" — spatial guidance that forces attention integration across multiple relevant regions rather than tunnel-visioning on one crop. The architecture play is elegant in its simplicity. They use procedurally generated synthetic scenes — think of it like arranging clip-art objects on a canvas with known identities and positions. This gives you unlimited annotation-free training data. The teacher (a frozen or EMA copy of the student MLLM) receives the question plus spatial guidance strings; the student receives only the image and question. The student's loss is computed against the teacher's token-level output distribution, not hard labels. This is standard on-policy distillation machinery (DAgger-family), but the innovation is in WHAT constitutes the privileged information and HOW cheaply it's produced. The ladder results are solid but not overwhelming. The headline number is a +3.23-point average gain across six real-world benchmarks (CVBench, V, ZoomBench, BLINK, HR-Bench, MME-RealWorld). The method is tested across multiple base models, which matters — it's not a one-architecture trick. On counting and document/chart understanding tasks, the gains are consistent. The honest caveat: the paper compares primarily against the base MLLMs and against Zoom-OPD, not against every possible fine-tuning or post-training strategy. The baselines are reasonable but not exhaustive. The integrity picture is mixed. The benchmarks used are community-standard and public, which is good. The synthetic-to-real transfer claim is the load-bearing wall of the paper, and it holds across multiple evaluation suites, which strengthens confidence. Code is indicated via a GitHub project page. However, there's no pre-registration, and the synthetic scene generation pipeline introduces a degree of freedom that could be tuned to flatter certain benchmarks. The fact that improvements hold across six diverse real-world benchmarks mitigates this concern substantially. The milestone question is about scale and generality. Today: +3.23 average points on perception benchmarks using synthetic scenes only. The next meaningful number would be demonstrating this approach on harder reasoning tasks (spatial reasoning, multi-step visual QA) at the 10+ point improvement level, and showing that procedurally generated scenes can be made complex enough to train genuinely hard perceptual skills — not just counting and chart reading. The successor experiment the authors did NOT run: training on procedurally generated scenes with compositional spatial relationships ("the cup is to the left of the book which is behind the lamp") to test whether the spatial guidance mechanism can teach relational reasoning, not just object localization. My read: this is being saved for the next paper — the infrastructure is clearly ready, and compositional spatial reasoning is the obvious next frontier.