Imagine you watched someone tie a knot once, then had to teach a thousand students who are all holding the rope differently. You wouldn't replay your exact hand motions — you'd describe the relationships: "this loop goes through that loop, then pull tight." That relational abstraction is exactly what Dex-One2Many does. Instead of forcing a robot hand to mimic the pixel-level trajectory of a single human demonstration, it extracts a sequential scene graph — a set of relational constraints between the hand, the tool, and the object at each stage — and uses those constraints to guide reinforcement learning in simulation. The committed claim: from a single human video, learn a dexterous manipulation policy that transfers zero-shot to a real multi-fingered robot hand and generalizes to object poses, goal poses, and grasps never seen in the demonstration. The authors report a 6.5% improvement over baselines on seen configurations, but the real headline is 71% improvement on unseen configurations. That gap tells the story — strict motion imitation works fine when the scene matches the demo, but collapses when anything shifts. Scene-graph guidance degrades gracefully. The architecture sits in the real-to-sim-to-real pipeline family, combining computer vision (hand/object tracking from video), graph-based task representation, and stage-wise RL with dense rewards. The scene graphs serve double duty: they generate diverse reset states for each training stage (sampling configurations that satisfy the relational constraints but aren't limited to the demonstrated pose), and they provide per-stage reward shaping that keeps the RL exploration tractable. This is clever engineering — high-dimensional dexterous manipulation with a five-fingered hand is notoriously hard for RL because the exploration space is vast, and most trajectories are garbage. By constraining relations rather than exact poses, the method threads the needle between imitation (too rigid) and pure RL (too unconstrained). On the ladder, the baselines include recent one-shot dexterous imitation methods. The 6.5% gain on in-distribution tasks is modest — these methods already work when the test configuration resembles the demo. The 71% gap on out-of-distribution scenarios is where the contribution lives. However, the evaluation is entirely self-benchmarked: five tasks in the authors' own simulation environment, with their own unseen-configuration protocol. There's no community-standard dexterous manipulation benchmark being used here, which makes independent comparison harder. Integrity is mixed. The sim-to-real transfer is demonstrated on a real Allegro hand across the five tasks, which is meaningful — real hardware doesn't lie. But the quantitative comparison numbers (6.5% and 71%) come from simulation, and the evaluation protocol (what counts as "unseen") is author-defined. Code and a project page are provided, which is a strong signal. No pre-registration, no independent replication yet, but that's standard for robotics. The milestone question for this line of work is: when does a single human video reliably produce a robot policy for a novel multi-stage task without any task-specific engineering? Dex-One2Many handles five tasks (tool use and object manipulation), but each requires the scene-graph extraction pipeline to correctly parse the video. The next concrete threshold is handling contact-rich tasks with deformable objects — cloth folding, cable routing — where relational graphs become harder to define. The authors didn't run these, likely because the scene-graph extraction from video becomes much noisier with deformables, and RL reward shaping for continuous deformation is an open problem. The broader field fight here is between imitation-first and RL-first approaches to dexterous manipulation. Pure imitation from human video is data-efficient but brittle. Pure RL is general but sample-hungry and often fails without curriculum design. This paper argues for a middle path: use the video to extract structure, then let RL discover the motor skills. It's a sensible bet, and the generalization numbers make the case. The question is whether the scene-graph abstraction scales to tasks where the relational structure itself is ambiguous or changes during execution.