You know those factory games — Factorio, Satisfactory — where at some point the bottleneck stops being individual machines and starts being the logistics network that routes materials between them? You can have the best smelter in the world, but if the belt layout is chaotic and nothing gets prioritized correctly, throughput collapses. This paper is about that transition point for AI agents: when the agent's task gets long enough, the biggest gains come not from better workers but from better routing. The committed claim: a structured controller-worker architecture for agentic inference, where a meta-reasoning controller explicitly decides which partial solutions to extend, when to start fresh, and when to stop — beating both production coding agents (Codex, Claude Code) and ablation baselines using the same workers and compute budget. On ProgramBench, the headline benchmark for long-horizon program reconstruction, meta-reasoning with GPT-5.5 hits 71.5% versus 58.0% for Codex. With Opus 4.8, it reaches 67.2% versus 65.5% for Claude Code. On abstract reasoning, multi-domain long-horizon reasoning, and proof generation benchmarks, gains of 3.6–4.2 points over a Direct Control Agent averaged across three frontier models. The architecture is conceptually clean. Workers do task-level computation. A controller maintains a compact state summary rather than replaying the full execution history — this is the key structural choice. Between steps, the controller consolidates what has been established, generates candidate next actions, estimates value-per-action under the remaining token budget, and dispatches work with context pulled from persistent memory. It is a decision loop that wraps around the agent's execution graph, not a modification of the underlying model. The approach builds an 'artifact graph' of partial solutions, enabling reuse of earlier work rather than starting from scratch each time. The ladder comparison is where this gets interesting. The baselines are not stale: Codex and Claude Code are current production systems, not two-year-old checkpoints. The Direct Control Agent ablation uses the same workers and the same compute budget, isolating the meta-reasoning contribution. The 13.5-point gap over Codex on ProgramBench is substantial. The 1.7-point gap over Claude Code with Opus 4.8 is modest but the scaling curves diverge — meta-reasoning keeps improving as budget increases where direct control plateaus, which is the more important structural finding than any single benchmark number. Integrity is mixed. The benchmarks span multiple domains (coding, abstract reasoning, proof generation), which reduces cherry-picking risk. The Direct Control Agent ablation is the right experiment — same compute, same workers, different control structure. But the evaluation is entirely same-team; there is no independent replication and no pre-registration. The authors acknowledge that meta-reasoning's overhead hurts at small budgets, which is an honest concession that raises confidence. The artifact-graph analysis (more reuse, higher coverage, nonuniform selection gains) provides mechanistic evidence beyond raw accuracy, though it is still self-evaluated. The milestone question is about scaling curves, not absolute numbers. The paper's core thesis is that structured control becomes MORE important as agents scale to longer runs. The demonstrated regime is roughly mid-tier: complex enough for the control overhead to pay off, but still within reach of production systems. The next concrete test is whether this architecture maintains its advantage at 10x the budget horizon — say, tasks requiring 100K+ token trajectories — where history replay becomes truly prohibitive and compact state summaries become essential. The obvious experiment not run: testing against the full Devin/SWE-bench pipeline on real software engineering tasks with actual code repositories, not program reconstruction. The authors chose ProgramBench, which tests long-horizon capability but in a controlled reconstruction setting. Real SWE tasks involve ambiguous specs, integration testing, and environment management. My read: the authors are either saving this for a follow-up with industrial partners (several are from Meta), or the overhead of the meta-reasoning loop is too expensive in wall-clock time for competitive SWE-bench results. Likely (c) — saving it.