Imagine you're teaching someone to solve sliding puzzles — not by writing step-by-step text instructions, but by handing them annotated photos of solved intermediate states with arrows showing what moved where. That's the core mechanism of ViSkill: instead of converting spatial game states into flattened text descriptions (which strips the geometry that matters), the system encodes successful interaction sequences as composite visual 'skill cards' — screenshots with overlaid action annotations that a vision-language model can directly consume. The committed claim: a visual-native skill library that co-evolves with PPO policy training outperforms text-centric skill approaches and raw VLM baselines on grid-world planning tasks, achieving 0.89 overall success rate (0.91 with cold-start seeding). The mechanism is a closed feedback loop — successful trajectories get distilled into visual skill cards, those cards shape both the inference prompt and the reward signal, and the improved policy generates better trajectories that produce better cards. It's skill accumulation and policy optimization reinforcing each other, which most prior skill-augmented agent work treats as separate pipelines. The benchmarks are Sokoban (box-pushing puzzle), FrozenLake (grid navigation), and PrimitiveSkill (action-composition tasks). These are well-understood environments, not frontier challenges. The paper compares against GPT-4o, Claude 3.5 Sonnet, Qwen2.5-VL-72B, and several open models as baselines, plus ablations stripping the visual skill component, the retrieval mechanism, and the cold-start module. The headline numbers are strong within this sandbox: ViSkill with cold-start hits 0.98 on FrozenLake, 0.80 on Sokoban, and 0.95 on PrimitiveSkill. Without cold-start, those drop to 0.95, 0.77, and 0.95 respectively. Architecturally, this is a retrieval-augmented VLM agent trained with PPO, where the retrieval targets are visual composites rather than text snippets. The skill cards themselves are constructed by stitching together key-frame screenshots from successful episodes with action-arrow overlays — essentially visual summaries that preserve spatial layout information that text linearization destroys. Retrieval uses CLIP-based similarity matching against the current observation. The cold-start mechanism seeds the library with expert demonstrations before online learning begins, which is practically useful but also means the 0.91 number includes a head start from human-provided examples. The integrity picture is mixed. The environments are deterministic grid-worlds, not stochastic real-world tasks. The paper doesn't compare against dedicated planning algorithms (A would trivially solve Sokoban and FrozenLake), which means the ladder is 'VLM agent vs. VLM agent,' not 'VLM agent vs. the best known solver for these tasks.' The ablation study is well-structured — removing visual skills, removing retrieval, removing the feedback loop each degrades performance — but there's no independent replication and the benchmarks are author-selected. Code is released. The real question this paper participates in is whether VLM agents should reason about spatial tasks through text or through vision. Most agentic frameworks (Voyager, DEPS, Cradle) convert observations to text, plan in text, and act from text. ViSkill argues that spatial reasoning should stay spatial — visual in, visual out. The 0.89 number is interesting not because Sokoban is hard (it isn't, for algorithms), but because it demonstrates that VLMs can learn to use visual reference material to improve their own spatial planning, which is a capability question with implications well beyond grid-worlds. The obvious successor experiment is scaling this to visually complex environments — Minecraft, web navigation, robotics simulation — where the visual skill cards would need to encode much richer spatial information and where retrieval similarity matching gets harder. The authors acknowledge this gap in their limitations. My read: they're saving it for the next paper, because the framework is designed to be environment-agnostic and the grid-world results establish the mechanism cleanly before tackling messier domains.