Imagine you're learning to draw rooms from memory. You sketch what you remember — colors, textures, furniture. Now someone hands you a ruler and says: also measure the depth of every object from where you're standing. You'd expect that second task to distract from the first. Instead, your color drawings get better, because knowing where things are in 3D forces you to understand their shapes, not just their appearances. That's the core mechanism of DepthWorld. The committed claim: a Stable Video Diffusion model that jointly predicts multi-view RGB and metric depth for robot manipulation produces better RGB predictions (+1.48 dB PSNR) than an identical model trained on RGB alone, while simultaneously yielding geometrically accurate depth — all without modifying the pretrained VAE. This is not a new architecture family; it's an engineering contribution showing that 3D supervision is a free lunch for video world models when you can get it cheaply enough. The data contribution may be more durable than the model. The authors build DROID-3D by running learned stereo depth (FoundationStereo) through a joint factor-graph optimization that pools every episode from the same physical robot to recover shared kinematic parameters and per-scene camera extrinsics. The result: dense metric depth and recalibrated multi-view extrinsics achieving <0.7 pixel reprojection error on 90% of episodes for external cameras. That calibration pipeline — not the diffusion model — is the piece other groups will reuse first. Architecturally, DepthWorld tiles spatial latents across views and modalities (RGB + depth) within the Stable Video Diffusion framework. The VAE stays frozen; depth is encoded into the same latent space via a separate channel. This is the conservative choice — minimal architectural novelty, maximum leverage of pretrained priors. The training budget is modest by foundation-model standards, and the spatial tiling trick is straightforward enough that replication should not be difficult. The integrity picture is mixed. Evaluation uses the DROID dataset's own held-out episodes with standard metrics (PSNR, SSIM, LPIPS for RGB; AbsRel, δ<1.25 for depth). The +1.48 dB PSNR gain is measured against the authors' own RGB-only baseline with identical compute, which is the right ablation but not an independent benchmark. No downstream policy evaluation is reported — the paper stops at perceptual quality and geometric accuracy, leaving the question of whether better world models actually produce better robot behavior unanswered. The field fight here is between RGB-only world models (UniSim, Genie, DIAMOND) that treat video prediction as a 2D problem, and approaches that insist 3D structure must be baked in for robotics to work. DepthWorld lands squarely on side B, arguing that 3D supervision doesn't just add a modality — it improves the original one. The +1.48 dB number is the evidence. Whether that improvement compounds into better downstream manipulation remains the open question. The obvious next experiment is closed-loop policy evaluation: train a manipulation policy using DepthWorld rollouts as synthetic experience, and measure task success rate against a policy trained on RGB-only rollouts or real data. The authors stop short of this. The honest read is (a) compute and experimental complexity — policy training on top of world-model training is a full second paper — combined with (c) saving it for the sequel. The CoRL 2026 acceptance suggests the community found the world-model contribution sufficient on its own, but the downstream payoff remains promissory.