Imagine you're an architect who builds a single scale model of a house, then discovers you can use that same model to plan the plumbing, the wiring, the furniture layout, and the demolition sequence — all without rebuilding. Most robot-learning systems build separate models for each of those jobs: one model to predict what the world will look like next, another to decide what action to take, another to figure out what happened. LeWAM collapses these into a single bidirectional transformer operating on a decoder-free JEPA latent space, trained end-to-end across four prediction modes: forward dynamics, backward dynamics, inverse dynamics, and policy prediction. The core claim is that this unified latent is both more informative and more disciplined than alternatives. Linear probes on LeWAM's latent read robot state and object state better than a forward-only JEPA world model (LeWM), while still ignoring visual distractors — something a reconstruction-based world-action model (think: autoencoder) fails to do. The latent is doing two things at once: capturing task-relevant information more faithfully, and filtering noise more aggressively. On the acting side, LeWAM's closed-loop policy matches a standalone flow-matching policy trained on the same encoder at the same model size — meaning you get world-model capabilities for free, without sacrificing control quality. This is a meaningful result: the worry with multitask latent spaces is always that cramming more objectives in will degrade any single one. LeWAM shows this tax doesn't have to be paid, at least at the scales tested. The planning contribution is the most architecturally interesting. When you do model-predictive control (MPC) with a world-action model, the standard approach is to sample candidate raw actions and roll them forward. LeWAM instead plans in the noise space of its diffusion-based policy head. The intuition: raw-action sampling can accidentally exploit inaccuracies in the dynamics model (finding trajectories that look good in simulation but are actually artifacts). Planning in noise space constrains the search to actions the policy already considers plausible, which acts as a regularizer. The authors report that this diffusion-steering-based MPC improves closed-loop performance over raw-action MPC. The architectural family is clearly JEPA (Joint-Embedding Predictive Architecture) — Yann LeCun's program for learning representations by predicting in latent space rather than pixel space, avoiding the generative-model tax of reconstructing observations. LeWAM extends this by making the JEPA bidirectional and multitask, and by grafting a diffusion policy head onto the latent. The transformer backbone handles sequence modeling across the four modes; the diffusion head handles the stochastic policy and enables the noise-space planning trick. Integrity-wise, the evaluation is simulation-based, likely on standard manipulation benchmarks, though the abstract doesn't name specific environments or provide absolute success-rate numbers. The comparisons are internal — LeWAM vs. LeWM vs. reconstruction-based WAM — rather than against a broad community leaderboard. Linear-probe diagnostics are a reasonable evaluation methodology, but they measure representation quality, not task performance at scale. The closed-loop claim of matching a standalone policy is load-bearing but needs independent confirmation. The paper is a solid architectural contribution at the intersection of world models and policy learning — the kind of work that advances a research program (JEPA for robotics) rather than solving a benchmark. The noise-space MPC idea is the most portable insight: it could transfer to any system combining a dynamics model with a diffusion-based policy. What's missing is scale — real-robot results, harder environments, and comparison against the strongest non-JEPA baselines would turn this from a plausible direction into a convincing one.