Imagine you're assembling IKEA furniture with only a YouTube tutorial. You don't rewatch the full 20-minute video before tightening each screw. You scrub to the relevant section, zoom in on the tricky joint, then zoom back out to see what comes next. That's the core mechanism here: RV-ICL turns a single demonstration video into a zoomable hierarchy the robot agent navigates on demand, rather than stuffing the whole video into a language-model prompt. The committed claim: a training-free, hierarchical video decomposition — built from sub-events like grasps and releases — lets an LLM-based robot agent retrieve exactly the visual detail it needs at exactly the right moment, beating flat video-in-context approaches by a meaningful margin on two established benchmarks. This is not a new VLA policy. It's a retrieval architecture layered on top of a frozen one (RPent), which means you can swap out the underlying policy without retraining RV-ICL. The hierarchy itself is the interesting engineering. The demonstration gets decomposed into levels — whole-task keyframes, phase-level segments, moment-level snippets, and short clips around contact events. These levels are exposed as read-only tools the agent can call. During planning, the agent reads the coarse levels. During execution, it re-enters the hierarchy and loads only the clip matching its current sub-goal. This solves the dual failure mode of prior video-in-context work: full videos bloat the context window and slow inference, while fixed keyframes lose the fine-grained contact detail that decides whether a grasp succeeds. The ladder position is clear but modest. Built on RPent, RV-ICL lifts success from 92.6% to 96.5% on LIBERO-PRO and from 86.7% to 95.8% on LIBERO-Plus. These are simulation benchmarks with 10 tasks each. The delta is real — about 4 and 9 percentage points respectively — but we're operating in the high-90s on sim tasks, not in real kitchens. The paper does not compare against other video-in-context methods head-to-head on the same benchmarks, which is an integrity gap worth noting. The integrity picture is mixed. LIBERO is a community benchmark, which is good. But validation is same-team simulation only; no real-robot experiments, no independent replication, and the paper doesn't report variance or confidence intervals in the abstract. The comparison baseline is RPent without RV-ICL — essentially an ablation against the system's own backbone, not against a competing video-retrieval method. Pre-registration is not mentioned. Code availability is not stated in the abstract. The milestone question is where this gets interesting. The real unlock isn't higher numbers on LIBERO — it's whether hierarchical video retrieval transfers to real hardware with real noise, real lighting, and real contact dynamics. The paper demonstrates that one demo is enough in simulation. The next concrete test: does one demo suffice on a physical manipulator doing a 5-step kitchen task with deformable objects? That's the number to watch — real-robot success rate with a single demonstration, no sim-to-real fine-tuning. The obvious experiment the authors didn't run is real-robot deployment. The honest read: this is almost certainly a compute/hardware access limitation, not a hidden negative result. The method is training-free by design, so there's no sim-to-real transfer gap in the learning sense — but the vision encoder and action decoder downstream may not handle real-world visual noise as gracefully as simulated scenes. The second missing experiment is comparison against other structured video retrieval baselines (e.g., VideoAgent, SuSIE) on identical benchmarks. That comparison would sharpen the claim from 'this helps' to 'this helps more than alternatives.'