Imagine you're a factory manager who replaced all your experienced floor supervisors with a giant rule book. Every supervisor used to watch a situation, interpret it on the fly, and make a call — expensive, slow, sometimes brilliant, sometimes wrong. Now imagine that instead, you wrote down every situation the factory could encounter as an explicit if-then rule, stored every piece of state in a ledger, and let a clerk follow the book. No interpretation, no salary for supervisors at runtime. The book gets thicker over time, and new tasks just add pages that reference existing ones. That's Code-Only-as-Policy. The committed claim: you can remove the vision-language model (VLM) and vision-language-action model (VLA) entirely from the robot's runtime control loop, replace them with pure Python code that explicitly tracks robot and environment state from camera images and proprioception, and achieve competitive performance on complex bimanual manipulation tasks. The authors frame this through the metaphor of an Embodied Turing Machine — the world is the tape, the code is the transition function, and the state is explicitly written down rather than hidden inside a neural network's activations. The ladder comparison is where things get interesting and also where you should squint. On RoboDojo's 42 bimanual tasks, COAP achieves 70.24% success rate. The paper compares against Agent Harnesses (VLM-in-the-loop systems like Agent-as-Policy and Harness VLA) and standalone VLAs, claiming advantages in cost, speed, controllability, and extensibility. But the comparison is structural rather than head-to-head on identical benchmarks with identical compute budgets. The 70.24% number is strong for a code-only system, but the paper is more convincing as an architectural argument than as a SOTA claim. Architecturally, COAP sits in the programmatic policy family — closer to classical robotics' state machines and behavior trees than to the end-to-end learning crowd. The key structural bet is that vision foundation models (used only for perception, not decision-making) can extract state accurately enough that all downstream reasoning can be hard-coded. The system uses a shared library of reusable code modules, and a coding agent (an LLM used at development time, not test time) iteratively builds and refines this library through what they call recursive self-improvement (RSI). The crucial distinction: the LLM writes the code offline; no model runs during execution. Integrity-wise, the validation is the authors' own benchmark (RoboDojo, 42 bimanual tasks) in simulation. There's no independent replication, no pre-registration, and the benchmark itself is relatively new. The 70.24% success rate is averaged across tasks, and with 42 tasks the variance could be hiding weak spots. The paper is honest about the upper bound being limited by state representation accuracy and code logic robustness, which is refreshing. But 'we tested our method on our benchmark' is a yellow flag for any paper making paradigm-level claims. The milestone question is concrete: COAP's ceiling is state representation fidelity. Today, vision foundation models extract state from RGB images with enough accuracy for 70% success on simulated bimanual tasks. The next number to watch is whether this transfers to real hardware with real perception noise — moving from sim-only to real-robot success rates above 50% on comparable task complexity would be the proof point. The gap is probably 1-2 years if someone puts serious engineering effort into it, but sim-to-real transfer has killed many promising approaches. The experiment they didn't run — and the one everyone will ask about — is real-robot deployment. The entire paper is in simulation. This is almost certainly (a): they didn't have the hardware setup or budget to run 42 bimanual tasks on real dual-arm robots. It's a reasonable omission for a first paper making an architectural argument, but it means the core claim ('code can replace models') is unproven where it matters most — in the messy, noisy, unstructured real world where state estimation breaks down.