Imagine you're trying to park a car using only the shadow it casts on the ground. You can tell roughly where the car is, but you've lost the door handles, the mirrors, the bumper contours — all the geometry that would let you thread a tight space. That's essentially what prior differentiable-rendering approaches to hand-eye calibration have been doing: rendering binary silhouettes of known objects, then optimizing the camera-to-robot transform by matching those silhouettes to observed images. DRHeC says: use the whole photograph. The committed claim is straightforward but consequential: by rendering full RGB images instead of binary masks and backpropagating color-plus-geometry gradients through the rendering pipeline, you get richer loss landscapes that avoid the local minima plaguing silhouette-only methods. The paper introduces a mask-guided image-to-image translation step — essentially a domain-bridging network that preserves color and geometric consistency between the synthetic render and the real camera image — so the comparison is apples-to-apples despite the sim-to-real gap. The ladder here is unusually clean. The named baseline is EasyHeC, the current state-of-the-art differentiable rendering method for markerless hand-eye calibration. On a UR5e robot performing real peg-in-hole tasks, DRHeC achieves 88.9% grasping success vs EasyHeC's 42.6%, and 57.4% insertion success vs 9.3%. Those are not incremental deltas — they're 46.3 and 48.1 percentage-point jumps on the same hardware, same objects, same task protocol. The paper also reports simulation experiments showing improved rotation and translation error metrics. Architecturally, this lives in the differentiable-rendering optimization family — not learned end-to-end from data, but using a physics-based renderer whose output is differentiable with respect to the extrinsic calibration parameters. The RGB extension adds a conditional image translation network (likely a lightweight U-Net or pix2pix variant with mask guidance) to bridge the domain gap. The optimization is gradient-based, iterating on the SE(3) transformation. The key compute property is that full RGB rendering is heavier per iteration than binary mask rendering, but the richer gradient signal means fewer iterations and fewer restarts from local minima. Integrity is solid but bounded. The real-robot experiments on a UR5e are genuine hardware validation — not just simulation. However, the task set is limited (grasping and insertion with specific objects), and there's no independent replication. The simulation experiments use their own setup rather than a community benchmark. The comparison to EasyHeC is direct and fair — same robot, same objects — but the paper doesn't compare against classical marker-based calibration (which would be the true upper bound) or against the latest learning-based methods outside the differentiable-rendering family. The milestone to watch is generalization. Right now this is validated on a single robot arm with a small set of objects. The question for the field is whether RGB differentiable rendering can become the default calibration pipeline — replacing both markers and learned keypoint methods — across diverse robots, grippers, and workspaces. A concrete next number: demonstrating comparable accuracy with 10+ object categories on 3+ robot platforms would make this a serious candidate for industrial adoption. The obvious experiment not run is a head-to-head against classical marker-based calibration (AprilTag, ChArUco) on the same tasks. The authors positioned this as a markerless method, so beating markers wasn't the stated goal — but the reader wants to know how close markerless now gets to the gold standard. My read: the authors know markers still win on raw calibration error, and including that comparison would complicate the narrative. They're saving the 'matches markers' claim for when the method matures further.