Imagine you're a restaurant manager who asks each server to declare their strategy before a shift: "I'll work section-by-section" or "I'll prioritize by table urgency." You discover two things. First, even when a server picks the right strategy, they abandon it mid-shift — wandering between sections, forgetting their own declared order. Second, when you force them to follow the strategy they picked by giving them a checklist that physically constrains their movement, performance nearly doubles. That's this paper. The committed claim: LLM agents exhibit a systematic Plan Declaration–Execution Gap — they can declare a planning mode (Predefined, Sequential, Hierarchical, or Search) but fail to faithfully execute it, and a deterministic routing mechanism that enforces structural fidelity dramatically closes this gap. The authors introduce Planning-as-Routing, where the LLM declares its planning mode and a hard-coded dispatcher channels execution to a pattern-specific executor rather than trusting the same LLM to both plan and follow through. The numbers are stark. On ALFWorld, generic Plan+ReAct achieves 0.48 task success; pattern-specific executors push that to 0.92. On SWE-bench Verified, the jump is 0.36 to 0.44. Across three benchmarks, only 22–45% of Plan+ReAct trajectories actually preserve the declared planning structure. The executors don't make the plans smarter — they make the agent do what it already said it would do. This is an execution paper, not a planning paper, and the distinction matters enormously. The architecture is refreshingly transparent. Four planning modes map to four executor templates. The LLM's only job is classification — pick a mode. A deterministic router dispatches accordingly. No learned components in the routing layer, no fine-tuning, no reinforcement learning. The method leans on the insight that LLMs are decent classifiers but unreliable self-monitors: they lose track of their own declared structure over long action horizons. By externalizing the structure, you get the LLM's language understanding without its execution drift. The integrity picture is solid but bounded. Four benchmarks (ALFWorld, SWE-bench Verified, WebArena, and a fourth) across three LLMs give reasonable coverage. The key finding — that pattern-specific executors enforce structure where Plan+ReAct doesn't — is measured via trajectory-level structural fidelity, not just final success, which is the right metric for this claim. However, all evaluation is same-team; no independent replication exists yet. The benchmarks are community-standard, which limits cherry-picking risk. The uncomfortable finding the authors are honest about: LLMs cannot reliably select the right mode for each task. Search wins on ALFWorld, Hierarchical wins on SWE-bench, and the best mode can vary across models within the same benchmark. Few-shot examples help sometimes but not consistently. This means the paper solves the easier half of the problem (execution fidelity) and leaves the harder half (mode selection) wide open. An oracle mode selector would likely push results further, but nobody has one. The successor experiment the authors didn't run — and the honest read is they're saving it for the next paper — is adaptive mode switching mid-task. Right now, mode is declared once at the start. Real tasks change character mid-stream: a sequential plan hits an unexpected state and needs search. The paper's framing practically demands this extension, and the authors gesture toward it in future work. The other obvious gap: scaling to more than four modes. Four is clean for a paper but arbitrary for real deployment.