You know that thing where you ask someone a question and they start answering confidently before they realize they have no idea? Vision-language models do this constantly. Ask about a red car in a photo that has no car, and the model will cheerfully describe its color, position, and state. This paper discovers that the model's internal routing decisions — the MoE gate — already encode 'I can't actually see that object,' and a trivially simple linear classifier on those signals catches the failure before the model opens its mouth. The committed claim: routing probabilities extracted from the target-object token in MoE VLMs (Qwen3-VL-30B-A3B and Gemma-4-26B-A4B) are sufficient to detect target-absence grounding failures with ROC-AUC above 0.99 on in-distribution data. An L2-regularized logistic regression — not a neural network, not fine-tuning — trained on these routing vectors achieves 0.9988 AUC on GQA-Inpaint for Qwen and 0.9956 for Gemma. When gated through a selective review prompt, end-to-end accuracy improves by +22.25% on GQA-Inpaint for Qwen. This is a detection-and-correction pipeline built from signals the model already computes, at effectively zero marginal cost. The architecture story matters here. MoE models route each token through a sparse subset of expert FFN blocks, and the routing probabilities form a signature of what the model 'thinks' about that token. The key finding is spatial and temporal: the signal localizes to the target-object token specifically (not the question tokens, not other visual tokens), and it emerges in early MoE layers, meaning the model makes its perceptual judgment early and the information propagates forward. The experts involved are partially substitutable — no single expert is the 'I see it' expert — which suggests the signal is distributed and robust rather than a brittle artifact. The integrity picture is mixed but honest. The primary benchmark, GQA-Inpaint, is a controlled dataset where objects are digitally removed from images to create ground-truth absence labels — a clean experimental setup. Cross-dataset transfer to OBER (a different absence benchmark) shows meaningful degradation: Qwen drops from 0.9988 to 0.8095 AUC, and Gemma from 0.9956 to 0.7781. The authors acknowledge this openly and attribute it to threshold shift requiring recalibration. They also note that false-positive reviews (prompting the model to double-check when the object is actually present) cause limited harm, which is a useful asymmetry for deployment. What's missing is any evaluation on truly in-the-wild data or adversarial inputs designed to fool the router. The ladder question — does this beat existing approaches? — is partially addressed. The paper positions itself as the 'first framework to leverage internal routing decisions' for this task, which is accurate in its specificity. Prior work used generated responses, hidden states, or uncertainty estimates. The routing-based approach has a real advantage: it's pre-generation, meaning you catch the failure before the model commits to a hallucinated answer. But the paper doesn't run head-to-head comparisons against hidden-state probing methods (like representation engineering or linear probes on residual streams) on the same benchmarks. This is the conspicuous missing experiment. The milestone trajectory is clear but early. Near-perfect in-distribution detection is achieved; the gap is cross-distribution robustness. Going from 0.81 AUC on OBER to 0.95+ on arbitrary out-of-distribution absence scenarios would make this production-ready for safety-critical VLM deployments. The authors suggest threshold recalibration as the path, but the deeper question is whether routing signatures generalize across visual domains (medical imaging, satellite imagery, autonomous driving) where absence detection actually matters. That's probably a 1-2 year horizon if the community picks this up. The obvious next experiment the authors didn't run: head-to-head comparison of routing-based detection against hidden-state linear probes and uncertainty-based detectors on identical benchmarks. My read is (c) — they're establishing the routing signal as a novel contribution and will benchmark against alternatives in a follow-up. The other missing piece is scaling: does this work on denser MoE architectures or models with different routing strategies (top-k vs. expert choice)? Likely a compute constraint for a three-author team.