You know how a good radiologist doesn't just stare at a scan? They pull up the patient's history, compare against reference images, check a database. They know when the image alone isn't enough. Most vision models don't do this — they get one look and commit. EviRover is the first perception agent explicitly trained to say "I need more information" and then go get it. The core claim is straightforward but consequential: visual perception should not be a one-shot prediction. When an image query requires fine-grained detail recognition or up-to-date world knowledge that the model's frozen parameters don't contain, the model should be able to interact — crop, zoom, search, retrieve — before committing to an answer. The authors call this "perception under insufficient evidence" and build a complete pipeline to address it: dedicated data, a new benchmark, and a two-stage training recipe. The data engineering is where the real work lives. Two pipelines generate EviRover-SFT-5K (5,000 supervised fine-tuning examples) and EviRover-RL-12K (12,000 reinforcement learning examples) that teach the model WHEN to seek evidence and HOW to use what it finds. EviLens, the human-verified benchmark, covers 688 instances across five perception categories designed to test exactly the cases where a single glance fails. This is the kind of infrastructure that makes the paper reproducible rather than anecdotal. The ladder position is strong for a 4B model. EviRover outperforms its backbone (likely InternVL2.5-4B or similar) by 30 points on average on EviLens. More impressively, it reaches performance comparable to advanced proprietary models — think GPT-4o-class systems — despite being roughly 10-50x smaller. The 15-point improvement on BrowseComp-VL and positive transfer to conventional perception and general multimodal benchmarks suggest the agentic capability isn't a narrow trick. The training recipe — supervised fine-tuning followed by agentic reinforcement learning — is the methodological contribution that other groups will likely adopt. SFT teaches the model the mechanics of evidence-seeking; RL teaches it the judgment of when to seek and when to commit. This two-stage approach echoes the SFT-then-RLHF pattern from language models but applies it to a fundamentally different problem: perceptual confidence calibration. The integrity picture is mostly clean. All code, models, and data are released — a full open-source stack. The benchmark is human-verified. Transfer to external benchmarks (WebEyes, BrowseComp-VL) provides some guard against overfitting to their own evaluation suite. The main weakness is the absence of independent replication and the fact that EviLens is their own benchmark, even if human-verified. The real question this paper poses to the field is whether agentic perception will become standard equipment in vision models the way tool use became standard in language models. If a 4B model can close the gap to proprietary systems simply by learning to gather evidence, the implication is that raw parameter count matters less than the ability to interact with the world. That's a significant architectural bet.