Imagine you're driving in an unfamiliar city with a friend who has Google Maps open. You don't need your friend to operate the steering wheel — you need them to say "turn left at the next intersection" while you handle the brake pedal, the lane changes, and the pedestrian who just stepped off the curb. Your friend thinks in blocks and intersections; you think in milliseconds and tire grip. WayFinder works exactly this way: a big multimodal language model sits offboard thinking in waypoints, while a tiny onboard controller handles the moment-to-moment flying. The committed claim: a zero-shot hierarchical VLA architecture that never fine-tunes the high-level MLLM for the robot embodiment, yet achieves up to 27.45% higher navigation success rates than flat low-level policies alone, across four AirSim environments of varying complexity. This is not a new model architecture per se — it's a systems-level decoupling argument. The insight is that you can treat the MLLM as a rarely-consulted strategic oracle rather than a continuous control loop, querying it only when the lightweight controller fails. On the ladder, the paper compares WayFinder against its own low-level baseline policies operating without high-level guidance. It tests three MLLM scales (GPT-4o, GPT-4o-mini, Gemini Flash) and reports that larger models produce better waypoints but the gap narrows in simpler environments. What's missing is comparison to any established VLA baseline like RT-2, SayCan, or other hierarchical navigation frameworks. The baselines are internal only, which limits how much the +27.45% number tells you about the state of the field. Architecturally, WayFinder sits in the hierarchical policy family — a high-level planner generating subgoals for a low-level executor. The high-level policy is a zero-shot MLLM processing linguistic context plus 2D state maps; the low-level policy runs onboard at high frequency using continuous sensor feedback. The key structural bet is asynchronous operation: the MLLM is queried only on navigation failure, not continuously. This is a compute-efficiency argument disguised as an architecture paper — it's exploiting the fact that cloud inference is expensive and latency-tolerant strategic decisions are rare relative to kinematic updates. Integrity is the weak link. All evaluation is in Microsoft AirSim — simulation only, no physical robot, no real-world transfer experiments. Four environments and three model scales give some variation, but the benchmarks are self-constructed. There's no pre-registration, no public leaderboard, and no code availability is mentioned. The paper is transparent about the simulation-only scope, but the reader should treat the 27.45% number as a simulation ceiling, not a deployment guarantee. The milestone question is concrete: WayFinder works in AirSim with four environments. The next number that matters is real-world transfer — can this hierarchy survive sensor noise, wind disturbances, and latency jitter on a physical drone? The gap between sim and real for drone navigation is well-documented; successful sim-to-real transfer with the same zero-shot MLLM and no fine-tuning would be a genuinely strong result. That's probably 1-2 years away given the IROS acceptance and the lab's trajectory. The obvious experiment not run: real-world deployment on a physical drone platform. The honest read is (a) — they ran out of time and resources for hardware experiments before the IROS deadline. Sim-to-real transfer for drone navigation is expensive and slow, and a clean simulation story is publishable at IROS. A secondary gap: comparison against existing hierarchical VLA systems (SayCan, Code-as-Policies, VoxPoser) that would contextualize the contribution against the broader field rather than just against their own low-level baselines.