Imagine you're solving a jigsaw puzzle, but every time you look away from the table, the pieces you've already placed vanish from your memory. You'd have to re-examine every piece each time you glanced back. Now imagine someone hands you a clipboard where you can pin snapshots of the table at every stage, flip back to any snapshot instantly, and rearrange which snapshots you're looking at as you reason about the next move. That's VISTA. The paper's committed claim: general-purpose multimodal models already have the reasoning chops for complex visual tasks — what they lack is a proper visual working memory. VISTA provides that memory as a harness around the model, not as a modification to the model itself, and the result is dramatic. On ARC-AGI-3, Claude Opus 5.0 goes from a Relative Human Action Efficiency (RHAE) of 40.68 to a perfect 100.00, completing all 25 public games with 57.4% fewer actions than first-time human players. The architecture is deceptively simple. VISTA sits between the model and the environment as a perception-and-memory layer. It captures raw visual observations — screenshots, grid states, whatever the environment produces — and stores them in a lossless visual memory bank. The model can actively retrieve past observations and reorganize its visual context window as it reasons. There's no learned memory module, no auxiliary network, no fine-tuning. It's a scaffolding pattern: take a frozen frontier model, give it the ability to look back at what it has seen, and let it decide what to re-examine. The key insight is that compressing observations into text or embeddings loses information that matters for spatial and visual reasoning. Keeping the raw pixels preserves the details the model needs. The ladder here is interesting. The paper's primary baseline is the same underlying model (Claude Opus 5.0) with minimal harnesses — essentially the model without VISTA's visual memory and retrieval. The jump from 40.68 to 100.00 RHAE on ARC-AGI-3 is the headline number, but the paper also tests across three additional benchmarks covering diverse visual games and puzzles, reporting substantial improvements across all of them. The baseline comparison is honest in one sense — same model, different scaffolding — but limited in another: we don't see comparisons against other agent harnesses like SWE-agent or Devin-style tool-use wrappers adapted for visual tasks. The paper positions itself as a general-purpose visual harness rather than a task-specific solver, which makes the benchmark breadth more important than depth on any single one. The integrity picture has strengths and gaps. ARC-AGI is a well-known community benchmark with clear evaluation criteria, and the 25 public games provide a concrete, reproducible test set. The additional three benchmarks add breadth. However, 25 games is a small sample — statistical variance matters at this scale. The paper is described as a tech report with an earlier version published as a blog post, which suggests the work is still evolving. No pre-registration, and the benchmark selection (ARC-AGI-3 specifically, plus three others) may reflect post-hoc choices. Independent replication on the public games would be straightforward and valuable. The milestone question is where this gets genuinely interesting. ARC-AGI has been a persistent thorn for AI systems — it was designed as a test of general fluid intelligence that resists memorization. A perfect RHAE score on the public set is a striking result, but the real test is the private set and the full 100-task evaluation. The trajectory to watch: can VISTA or its successors hit >90% on the full private ARC-AGI-3 set? That would be a much stronger signal than 25/25 on the public games. The obvious experiment the authors did not run: VISTA on the full private ARC-AGI-3 evaluation (not just 25 public games), and VISTA with models other than Claude Opus 5.0 — particularly open-weight models where the community could replicate and extend. The honest read is probably (c): they're saving model-breadth experiments for follow-up work, and the private evaluation may require coordination with the ARC Prize Foundation. The blog-to-paper pipeline and the author list (including Kaiming He) suggest this is positioned as a foundational framework paper, with the headline ARC-AGI result as proof of concept rather than the endpoint.