Imagine you're directing a play, but you can only give instructions by pointing at spots on the stage floor. You say "move from here to there," but the actor doesn't know if you mean walk forward, climb stairs, or lean sideways — because you flattened 3D reality into a 2D map. That's the core problem with current controllable video generation: a 2D trajectory on screen could mean the object moved left, the camera panned right, or both happened simultaneously. GenCine's fix is to give the director a proper 3D stage. The committed claim: GenCine decouples camera motion from foreground object motion in 3D space, encodes both into color-coded guidance maps, and trains a lightweight branch on a pretrained Wan video model to follow those controls — producing geometrically consistent video even when camera and objects move simultaneously. This is not the first system to offer controllable video generation, but it is among the first to resolve the camera-vs-object ambiguity by working in a shared 3D world coordinate system rather than screen space. Architecturally, GenCine sits in the guidance-conditioned diffusion family. It takes a single image, lifts it into a 3D scaffold (depth estimation plus segmentation), and lets users place local 3D motion handles on foreground regions. Multiple handles can drive different parts of the same subject — a piecewise-rigid approximation to non-rigid motion that sidesteps physics simulation entirely. The handles' 3D positions are projected into per-frame guidance maps where each handle gets a fixed color encoding its world-coordinate position. These maps are fed to a guidance branch and LoRA adapters grafted onto Wan 2.1, a pretrained video diffusion model. Training data comes from two sources: real videos where controls are recovered via monocular depth and tracking, and synthetic videos with ground-truth geometry. The ladder question is instructive. Prior work like MotionCtrl, CameraCtrl, and DragAnything operate in 2D screen space — they can move things around in the frame but cannot distinguish camera motion from object motion. Direct3D and other recent camera-control methods handle the camera but not foreground objects. GenCine's contribution is the joint control: camera path plus per-handle 3D object trajectories, resolved in a common coordinate frame. The authors show qualitative improvements in geometric consistency under viewpoint changes and more faithful adherence to intended 3D motion, though the evaluation leans heavily on qualitative comparisons and user studies rather than hard quantitative benchmarks against all named competitors on a shared metric. Integrity-wise, the validation is a mix. Training uses both real video (recovered controls) and synthetic video (ground-truth geometry), which is a reasonable two-source strategy. But the evaluation is primarily qualitative — visual comparisons, ablations, and a user study. There is no single community benchmark for this task because the task itself (joint 3D camera + object control) is new enough that no standard evaluation protocol exists yet. This makes the results hard to compare rigorously. The synthetic-to-real transfer is plausible but not independently validated. Code and model availability are not explicitly confirmed in the abstract. The milestone that matters is whether this 3D-scaffold-plus-guidance-map paradigm can scale to complex multi-object scenes with genuine non-rigid deformation — cloth, fluid, hair — without falling back to physics simulators. Right now the piecewise-rigid approximation handles simple articulated motion (a person walking, an arm swinging). The gap to full non-rigid control is significant. A concrete next marker: controllable generation of a scene with 5+ independently moving non-rigid objects maintaining 3D consistency across 10+ seconds of video at 720p or above. The obvious experiment not run is quantitative evaluation on a standardized benchmark with full numerical comparisons against the strongest recent competitors (e.g., CogVideoX-based controls, Kling's camera system). The honest read: the task is new enough that no such benchmark exists, and building one is a paper-sized effort. But the absence means we're taking geometric consistency on faith from cherry-picked examples — a common and understandable gap at the "new task definition" stage, but one that needs closing fast.