Imagine you're trying to judge the distance to a lamppost from a moving car. You could train a neural network on millions of photos of lampposts at known distances — or you could just look at the lamppost from two slightly different positions, note how much it shifted in your view, and do the trigonometry. That second approach is what this paper does for drones, and it turns out the trigonometry wins when the neural network has never seen drone altitudes before. The core claim: a purely geometric, training-free method called EpiTransfer synthesizes a virtual stereo pair from two temporally separated monocular drone images and known camera poses, then runs standard stereo triangulation. No learned weights, no domain-specific training data. The trick is choosing an arbitrary virtual baseline for the synthesized stereo rig, sidestepping the geometric degeneracies (near-zero parallax, forward-dominant motion) that plague naive two-view triangulation from a drone's typical flight path. The method sits in the classical multi-view geometry family — specifically epipolar transfer via the trifocal tensor, a technique dating back to Hartley and Zisserman's textbook. What's new is not the math but the application context: using camera-pose estimates from a drone's onboard IMU/SLAM to construct the virtual baseline on the fly, tuned to the scene geometry. The computational backbone is sparse feature matching (SuperPoint + LightGlue), not dense prediction, which keeps the pipeline lightweight but means output is sparse point clouds, not dense depth maps. Against the ladder: indoor LiDAR ground truth gives EpiTransfer an AbsRel of 0.092 and δ<1.25 of 0.940. Direct two-view triangulation edges it at 0.073 AbsRel but recovers fewer valid points in challenging geometry. The real story is the comparison to learned baselines: ZoeDepth hits 0.225 AbsRel, Depth Anything V2 crashes to 0.570. Those numbers aren't a fair fight — neither learned model was fine-tuned for aerial imagery — but that's precisely the point. When you fly a drone somewhere the training set has never been, the geometry-only approach doesn't degrade. Outdoor validation extends to approximately 90 meters range, but the paper is notably thinner on outdoor quantitative numbers. Integrity is a mixed bag. The LiDAR ground truth for indoor scenes is solid, and the comparison to direct triangulation provides a meaningful geometric baseline. But the learned baselines (ZoeDepth, Depth Anything V2) are used off-the-shelf without any domain adaptation — beating an out-of-distribution neural network with in-distribution geometry is less impressive than beating a fine-tuned one. The paper acknowledges this, but the headline numbers could mislead a skimmer. No code release is mentioned, and the outdoor evaluation lacks the quantitative rigor of the indoor experiments. The milestone question is about density and range. Sparse depth at 90 meters is useful for obstacle avoidance but insufficient for dense reconstruction. The obvious next step — and the experiment conspicuously absent — is integrating a dense matcher (or a learned depth prior for densification) while keeping the training-free geometric backbone. My read: the authors are likely saving this for a follow-up, since the sparse-to-dense bridge is a natural extension and wouldn't require new data collection. This paper won't reshape computer vision, but it delivers a sharp reminder: when domain shift is your primary enemy, classical geometry with modern feature matchers can outperform deep learning's billion-parameter depth estimators. For robotics teams deploying drones in novel environments — disaster response, infrastructure inspection, exploration — that's a practical and immediate insight.