Imagine you're directing someone to parallel park over the phone. You could describe every pixel of the scene — the color of the curb, the shadow under the SUV, the texture of the asphalt — or you could say 'your front bumper is two feet from the white car's rear corner, and the curb is eighteen inches to your right.' SkeleWAM bets that robot manipulation works the same way: the geometry that matters for control fits in a sparse skeleton of 3D points, and everything else — textures, lighting, background clutter — is noise the policy has to learn to ignore. The committed claim: representing a manipulation scene as a sparse 3D skeleton of robot joints, object centers, and interaction points provides a more effective state space for world-action modeling than video or learned visual latents, achieving 85.9% success on LIBERO-Plus with 57.1M parameters — beating Cosmos-Policy (82.2%) at roughly 18× fewer parameters. This is not the first paper to argue that geometric representations beat pixels for manipulation, but it is among the first to build a full world-action model (joint action generation and future state prediction) on top of such a minimal skeleton. The architecture sits in the world-action model (WAM) family — methods that interleave action generation with future-state prediction so the model learns dynamics, not just reactive control. Where prior WAMs like Cosmos-Policy predict future video frames or latent visual states, SkeleWAM predicts future skeletons — the next positions of those same sparse 3D points. The skeleton is constructed online from RGB-D observations and robot proprioception, meaning no offline 3D reconstruction or privileged simulation state is needed. A Medoid Action Consensus (MAC) strategy filters stochastic action samples at inference, selecting the medoid of a sampled set to reduce variance without requiring ensembles. The ladder comparison is honest but narrow. LIBERO-Plus is the headline benchmark, and SkeleWAM's 85.9% overall success rate tops Cosmos-Policy's 82.2% and several other recent WAMs. The 57.1M parameter count against Cosmos-Policy's roughly 1B+ is the real headline — it's not just better, it's radically smaller. But LIBERO-Plus is a simulated tabletop benchmark. The paper does not test on real hardware, and it does not compare against non-WAM manipulation baselines like RT-2 or Octo on their own turf. The win is clean within its benchmark, but the generalization question is wide open. Integrity is mixed. The benchmark is a recognized community suite (LIBERO), and the paper compares against multiple recent WAMs on the same tasks, which is good practice. But all evaluation is in simulation — no real-robot transfer, no domain-randomization stress tests, no out-of-distribution object categories. The skeleton construction relies on RGB-D depth quality and object detection/segmentation, failure modes the paper does not deeply probe. Code and project page are available, which helps replication. The milestone to watch is real-robot transfer with noisy depth sensors. Simulated RGB-D is clean; real Intel RealSense or Kinect depth is not. If SkeleWAM's skeleton construction degrades gracefully under real sensor noise — say, maintaining >70% success on a 10-task real-robot suite with commodity depth cameras — that would validate the core thesis outside the simulation sandbox. The gap is probably 1-2 papers away, not a hardware generation. The obvious experiment the authors did not run is real-robot deployment. The honest read: this is likely (a) — real-robot setups are expensive and slow, and the simulation results are strong enough to publish. There's no sign they're hiding bad results; the project page suggests ongoing work. A secondary missing experiment is scaling the skeleton to cluttered multi-object scenes with occlusion, where sparse keypoints may collide or disappear. That's probably (c) — saved for the next paper.