Imagine you're searching for a lost dog in a sprawling park. You have a friend in a helicopter who can see the whole park but can't read the collar tag, and you're on the ground where you can read every sign but can only see 50 meters ahead. The helicopter spots a brown blur near the pond and radios you directions, but by the time you get there, the dog has moved — and the helicopter's earlier observation is now stale. This paper is about building the shared notebook and planning protocol that makes that relay work for robots. The committed claim: AirGroundVLN is the first large-scale benchmark specifically designed for goal-oriented air-ground collaborative vision-and-language navigation, paired with a trainable framework (AG-CoNAV) that addresses the two core problems of cross-platform memory and asymmetric planning. Prior VLN benchmarks (R2R, REVERIE, SOON, Room-to-Room) are almost exclusively single-agent and ground-only. The handful of aerial VLN works (AerialVLN, CityNav) treat the drone alone. Nobody had built a systematic testbed where a drone and a ground robot must cooperate on language-described targets with no route prescribed. The benchmark itself is substantial: 10,281 navigation episodes, 955 distinct target instances, 19 Unreal Engine environments with seen/unseen splits and an aerial-visibility protocol that controls whether the drone can actually see the target. This last detail is critical — it lets you separately measure whether the system degrades gracefully when the aerial advantage disappears. The environments span indoor, outdoor, and mixed scenes, which is a meaningful step beyond the single-building or single-city setups that dominate VLN. The AG-CoNAV framework has two load-bearing modules. Spatiotemporally Anchored Collaborative Memory (SACM) maintains a shared spatial memory that fuses aerial panoramas and ground observations over time, anchored to spatial coordinates so that older observations don't just vanish when new ones arrive. Aerial-Guided Regional-to-Local Planning (AGRLP) uses the drone's broad view to identify a promising region and then hands off to the ground agent for fine-grained local navigation. The architectural family is attention-based cross-modal fusion (visual-language grounding) combined with a hierarchical planner — squarely in the VLN Transformer lineage (VLN-BERT, HAMT, BEVBert) but extended to two heterogeneous agents. Integrity is mixed. The authors evaluate on their own benchmark — which is inevitable for a first benchmark paper — but they compare AG-CoNAV against several adapted baselines (Seq2Seq, CMA, HAMT, BEVBert) plus ablations removing SACM and AGRLP individually. The seen/unseen splits and aerial-visibility protocol are well-designed controls. However, all evaluation is in simulation (Unreal Engine), no code availability is confirmed in the abstract, and there is no independent replication. The benchmark design itself is the main contribution; AG-CoNAV is framed as a reference framework, not a competition winner. The milestone question is whether air-ground collaborative VLN can move from simulated Unreal environments to real-world deployment. Today: 19 simulated environments. The next meaningful number is a sim-to-real transfer demonstration on even 2-3 physical environments with real drone-robot hardware. That probably requires solving domain gap plus real-time communication latency — likely 2-4 years out for a convincing demo. The obvious experiment not run: real-world transfer. The authors built everything in Unreal Engine, which gives control but leaves the sim-to-real gap entirely unaddressed. The honest read is (a) — real hardware experiments with a drone and ground robot are expensive, slow, and require institutional resources beyond most academic labs. They're almost certainly planning it, but the benchmark needed to come first to justify the hardware investment.