Imagine you're training a new taxi driver. During lessons, you sit in the passenger seat with a GPS, satellite imagery, and a bird's-eye drone feed, whispering the correct route whenever the trainee makes a wrong turn. But on their first solo shift, you take all of that away — they only have their eyes and whatever spatial intuition your coaching burned in. That's GPD: geometry-privileged distillation. The committed claim is that on-policy self-distillation with routed 3D privilege — depth maps, semantic labels, and bird's-eye-view renders injected as compact text into the teacher only — beats both pure reinforcement (GRPO) and answer-only privileged distillation on spatial reasoning benchmarks, while the deployed student remains a standard RGB-input VLM with zero inference overhead. The paper reports 57.1 on VSI-Bench and a 37.6 average across four additional spatial benchmarks (MindCube, SPARBench, MMSI-Bench, ViewSpatial) on a 4B-parameter backbone. The ladder here matters. The baselines are GRPO (reward-only RL) and answer-privileged OPSD (which gives the teacher the reference answer but no geometry). GPD beats both. On VSI-Bench, the delta over GRPO is meaningful — geometry privilege adds what outcome supervision alone cannot: perceptual correction at the trajectory level. However, the comparisons are entirely within the authors' own training framework on one backbone size (4B). There is no comparison to larger proprietary VLMs or to methods that inject 3D features at inference time and pay the latency cost. The paper is honest about this scope but it does mean the ladder is internal, not field-wide. Architecturally, GPD sits in the GRPO family (group relative policy optimization, itself a variant of PPO-style RL for language models) augmented with a privileged-information KL divergence term. The key structural choice is question-conditioned routing: a lightweight classifier decides which 3D cue (depth, semantic, BEV) to inject for each question, rather than dumping all geometry in at once. This is a genuine design insight — full-context injection actually hurts, because irrelevant 3D evidence adds noise. The privilege is rendered as text tokens, not as visual features, which keeps the teacher's architecture identical to the student's and avoids multimodal fusion complexity. Integrity is mixed. The benchmarks — VSI-Bench, MindCube, SPARBench, MMSI-Bench, ViewSpatial — are community-established spatial reasoning benchmarks, which is good. But the ablations are all internal: same team, same codebase, same backbone. There's no pre-registration, and the benchmark selection could reflect post-hoc curation (five benchmarks were chosen; we don't know if others were tried). Code is released on GitHub, which is a strong signal — it invites replication. But no independent group has reproduced these numbers yet. The milestone question for spatial VLMs is clear: when does a pure-RGB model match a model that has actual depth sensors at inference time? GPD narrows this gap but doesn't close it. The next concrete number to watch is whether GPD's approach scales to 7B–13B backbones and whether the gains compound or plateau. If a 13B GPD model matches an inference-time-3D model on VSI-Bench (currently ~60+ for augmented models), that's the threshold where this technique becomes standard practice rather than a research demonstration. The obvious experiment not run is scaling beyond 4B parameters. The paper demonstrates on a single backbone size. Running at 7B or 13B would show whether privileged distillation's gains are additive with scale or get absorbed by the larger model's inherent capacity. My honest read: (a) compute budget — training multiple backbone sizes with RL is expensive, and this is a university lab (ZJU), not a frontier lab with thousands of GPUs. The 4B result is the proof of concept; scaling is the next paper.