Imagine you commute to work every day. The first time, you use GPS turn-by-turn — expensive, slow, constantly recalculating. By the tenth time you know the route cold and drive it from muscle memory; you only glance at the GPS when construction forces a detour. PhoneCLI applies this exact logic to mobile AI agents: compile the routes once, replay them deterministically, and only invoke the expensive vision-language model when the compiled path breaks. The committed claim: you can extract a usable command-line interface from any mobile app's GUI — without internal APIs, runtime hooks, or model fine-tuning — by exploring the app offline, mapping its screens and transitions, and converting each reachable screen into a deterministic replay sequence. The agent then picks commands from this menu instead of reasoning through every tap. This is not a new agent architecture; it's a compilation layer that sits in front of any existing VLM agent and makes the routine part free. Architecturally, PhoneCLI belongs to the GUI-grounding agent family (AppAgent, CogAgent, M3A) but introduces an offline exploration-and-compilation phase that produces a semantic graph of an app's navigation structure. Each node is a screen with its interactive elements annotated; each edge is a deterministic action sequence. Online, the agent matches a task to a command, runs a pre-execution verification step, and replays. The key structural bet is that most mobile interaction is navigation — getting to the right screen — and navigation is static enough to precompute. When it isn't (dynamic content, novel screens, failed replays), the system falls back to the full VLM agent, meaning compilation can only help and never hurt the success rate. The ladder comparison uses AndroidLab and AndroidWorld, two community benchmarks for mobile agents. On AndroidLab, PhoneCLI improves task success rate over the base VLM agent while reducing both step count and token consumption. It also transfers to AndroidWorld's official M3A agent with consistent efficiency gains. The paper does not report exact percentage improvements in the abstract, and the baseline is the embedded VLM agent itself rather than a competing compilation approach — because no direct competitor exists in this exact niche. The honest framing is: this is a new capability layer, not a SOTA race on a fixed leaderboard. Integrity is reasonable but not airtight. AndroidLab and AndroidWorld are recognized community benchmarks, not cherry-picked toy tasks. The zero-training claim is structurally verifiable — the method literally doesn't touch model weights. However, the exploration phase presumably has parameters (exploration depth, screen similarity thresholds) that could be tuned to the benchmark apps. No pre-registration, no code release mentioned in the abstract, and no independent replication. The 'compilation can only help' guarantee is the strongest structural claim: since failures fall back to the unmodified VLM agent, the worst case is identical performance at no extra inference cost. The milestone question for this line of work is coverage: what fraction of real-world app interactions can be compiled away? PhoneCLI compiles navigation, which the authors argue is the majority of mobile agent activity. The next concrete target is pushing compiled coverage past routine navigation into conditional flows (forms, search results, dynamic lists) — the territory where replay sequences break. If someone demonstrates 80%+ task coverage with compiled commands across the top 50 Android apps, this approach graduates from research demo to production infrastructure. The obvious experiment not run: scaling the exploration to hundreds of apps and measuring how the compiled command set degrades over app updates. Apps ship new UI versions constantly; a compiled map that rots in two weeks has limited practical value. The honest read is (a) — this requires sustained engineering effort and real-device farms that a research team may not have, not a result they're hiding.