Imagine you're watching a cooking competition. You know the contestant is making risotto the moment you see arborio rice hit the pan — you don't need to wait for plating. Your brain didn't learn a separate "have I seen enough" rule; the same visual processing that identifies risotto also knows when the evidence is sufficient. This paper asks whether video-language models work the same way, and the answer is yes — and it's linearly decodable from their frozen weights. The committed claim: frozen VideoLLMs already compute an internal "evidence readiness" signal that is question-conditioned, linearly readable, and independent of answer correctness. No fine-tuning, no architectural modification, no learned trigger module. The signal is just sitting there in the activations, waiting for a linear probe to extract it. Across seven models evaluated on byte-identical video inputs, the probe achieves AUROC between 0.733 and 0.905 under the strictest sampling condition — where a naive clock baseline is near chance. That's the headline number, and it's robust. What makes this more than a probe curiosity is the question-conditioning result. On identical video windows, changing only the text question reverses the readiness readout on 66.1% of paired samples, while every question-blind control sits at chance by construction. This rules out the boring explanation that the model is just tracking general "visual complexity" or "temporal completeness." It's tracking whether THIS question has enough evidence in THIS window. Even more striking: the signal persists when the model gets the answer wrong (AUROC 0.722), meaning readiness is not just a confidence proxy — it's encoding something closer to "the evidence is present" rather than "I'm sure of my answer." The ladder here is informative. The paper compares readiness probes against uncertainty estimators (entropy, token probability, semantic uncertainty) and their supervised combination on latency-matched answer selection. Readiness wins. It also correlates more closely with independent human judgments than confidence scores do. Existing streaming trigger modules — which ARE trained — turn out to be approximately orthogonal to the readiness signal and decode it far less accurately than the simple linear probe. This is a genuine embarrassment for the trigger-learning paradigm: the feature they're trying to learn is already there, and their learned triggers aren't even using it. Architecturally, this is a probing study on the transformer-based VideoLLM family (the paper tests seven models sharing byte-identical evaluation pipelines). The method is a linear probe on frozen intermediate activations — computationally negligible. The key structural insight is that readiness lives in a different subspace than the model's own streaming triggers, which suggests current trigger architectures are solving an optimization problem in the wrong direction. The practical payoff is Readiness Gating: a policy that simply waits until the readiness probe fires before committing an answer, yielding up to +9.75 percentage points accuracy at matched video duration. Integrity is solid for a probing study. The byte-identical evaluation across seven models is a strong design choice — it eliminates confounds from different preprocessing pipelines. The cross-benchmark generalization test (probe trained without any footage from the target benchmark family still reads that family) guards against overfitting to specific visual patterns. The 26-configuration sweep showing gain tracks available headroom is honest about when the method helps and when it doesn't. The main gap: no independent replication, and the probe-vs-trigger comparison is run on the trigger's own base model, which is fair but could benefit from third-party validation. The milestone question is where this gets interesting for practitioners. Right now, readiness gating gives +9.75 pp at best, but the gain is bounded by the accuracy headroom available in the task. The next concrete number to watch: can readiness gating close more than half the gap between streaming and full-video accuracy on hard benchmarks (ActivityNet-QA, NExT-QA temporal splits)? If a follow-up shows readiness gating at 50%+ headroom closure on temporal reasoning tasks, that's the signal that this replaces learned triggers entirely. The obvious experiment NOT run: training a trigger module that is explicitly initialized or regularized to align with the readiness subspace. The honest read is (c) — they're saving it for the next paper, because the current narrative is cleaner as a pure probing discovery.