Imagine you're a line cook who has mastered dicing onions, and now someone asks you to dice carrots. You don't learn knife work from scratch — you retrieve the dicing pattern from muscle memory and adapt it for the new vegetable's density. That's the core mechanism of this paper: building an explicit memory bank of physical manipulation skills that a robot can look up and adapt, rather than re-deriving every action sequence from pixels alone. The committed claim is that World Action Models (WAMs) — systems that jointly predict visual futures and generate actions — suffer from an inability to reuse prior action experience across tasks. The authors propose an Action Experience Dictionary (AED), a learned lookup table of action embeddings compiled from historical trajectories. When a new task arrives, the model retrieves relevant embeddings from the dictionary, conditions them on the current visual scene via cross-attention, and prepends them to noisy action tokens for diffusion-based prediction. This is not the first skill-reuse mechanism in robotics, but the dictionary-as-shared-embedding framing within the WAM paradigm is a distinct architectural choice. Architecturally, the system sits in the vision-conditioned diffusion policy family. A pretrained action tokenizer quantizes historical trajectories into discrete tokens that index into the AED. Cross-attention fuses these retrieved embeddings with visual features, and the combined representation seeds a denoising diffusion process that outputs action sequences. The second contribution — a motion-aware transition loss — supervises the model to predict visual feature changes across random temporal intervals, steering attention away from static background clutter toward task-relevant motion. This is a lightweight but smart regularization trick. On the ladder, the authors benchmark against several recent WAMs and policy-learning methods on simulation benchmarks (the abstract mentions simulation benchmarks and real-world cross-embodiment settings without naming specific benchmark suites or providing delta numbers). The absence of named baselines with specific numeric margins in the abstract is a notable gap — we're told the approach "verifies effectiveness" but not by how much against which named competitor. The real-world cross-embodiment experiments are the strongest credibility signal, since sim-to-real and cross-embodiment transfer are where most methods collapse. Integrity is mixed. The simulation benchmarks provide controlled comparison, but without knowing the specific suites (CALVIN? RLBench? MetaWorld?) and whether they were selected before or after method development, the usual cherry-picking risk applies. Code is available via the anonymous GitHub project page, which is a positive signal for reproducibility. No pre-registration, no independent replication — standard for the subfield, but worth noting. The cross-embodiment real-world experiments add genuine validation weight. The milestone question for this line of work is generalization breadth: how many distinct manipulation skills can a single AED store before retrieval degrades, and at what dictionary size does the approach match or exceed specialist policies trained per-task? The paper doesn't state these thresholds explicitly. The trajectory to watch is whether dictionary-based skill reuse scales — from tens of tasks in controlled sim to hundreds in diverse real-world kitchens and warehouses. The obvious unrun experiment is scaling the AED to a much larger, more diverse action corpus — say, thousands of distinct manipulation trajectories across dozens of embodiments — and measuring retrieval degradation and interference. The honest read: this is likely (a) compute-limited and (c) being saved for the next paper. A second gap is ablating the motion-aware transition loss against other background-suppression techniques (e.g., object-centric representations). The authors introduced it as complementary but didn't stress-test alternatives.