Imagine you're debugging a factory assembly line, but instead of watching every widget move through every machine, you discover that every intermediate state of the line can be described by a compact checklist — no matter how complex the machinery looks. You never need to track exponentially many possibilities because the checklist always stays short. That's what this paper proves for a critical class of quantum error-correction circuits: every intermediate state fits inside a tractable mathematical structure, even when noise is present. The committed claim: a broad class of non-Clifford quantum error-correction circuits — including magic state distillation, cultivation, code switching, gauge fixing, transversal non-Clifford gates with syndrome extraction, and diagonal magic state injection — can be simulated exactly in polynomial time on a classical computer. This is not an approximation or a heuristic. The authors prove that every intermediate state in these circuits is a "third-order phase-polynomial state," a structure they show is equivalent to stabilizer states under a new formalism they call diagonal-Clifford-and-Pauli (DCP) stabilizers. The stabilizer formalism is the workhorse of efficient classical simulation in quantum computing — Gottesman-Knill showed Clifford circuits are poly-time simulable decades ago. This paper extends that boundary outward to cover the non-Clifford operations that are essential for fault-tolerant quantum computing. The ladder here matters enormously. The existing tools for simulating non-Clifford circuits — like those based on stabilizer decomposition or direct state-vector simulation — scale exponentially in the number of non-Clifford (T) gates or qubits involved. The authors benchmark their open-source tool, merlin, against these simulators on real distillation, cultivation, and code-switching circuits. On Bravyi-Haah distillation, merlin demonstrates improved runtime and memory scaling as logical outputs grow. On a code-switching circuit, merlin succeeds where all other tested simulators fail entirely. The key insight is architectural: by recognizing that DCP stabilizers form a group, the authors can update state representations through matrix operations over finite fields rather than tracking exponential superpositions. The integrity story is strong for a theory paper. The core result is a mathematical proof — not a simulation validating itself. The claim that third-order phase-polynomial states are exactly DCP stabilizer states is a structural theorem, not an empirical observation. The benchmarks against other simulators serve as practical validation, and the code is open-source, which means independent verification is possible immediately. The choice of benchmark circuits (distillation, cultivation, code switching) covers the major non-Clifford QEC techniques in active use, not cherry-picked toy problems. What this unlocks is less about simulating quantum computers for fun and more about accelerating the design cycle for fault-tolerant quantum error correction. Today, designing and verifying QEC protocols is bottlenecked by the inability to classically simulate the non-Clifford components at scale. If you can simulate these circuits in polynomial time, you can iterate on protocol design orders of magnitude faster — test new distillation schemes, explore code-switching strategies, optimize magic state factories — all on classical hardware. This is the equivalent of giving chip designers a fast SPICE simulator instead of making them build every prototype in silicon. The obvious successor experiment: scaling merlin to the full magic state factory designs proposed for surface-code architectures (hundreds to thousands of logical qubits with deep non-Clifford layers). The authors demonstrate the formalism works on the circuits they tested, but the real stress test is whether polynomial-time scaling holds in practice — not just in theory — when applied to production-scale QEC designs that groups like Google and IBM are actively developing. The honest read: they likely ran out of compute for the largest circuits and are saving the industrial-scale benchmarks for a follow-up, possibly in collaboration with a hardware team that has specific factory designs to verify.