Imagine you're building a cathedral out of dominoes. Each room depends on the rooms beneath it, and every domino has to lean exactly right. Now imagine you discover that one domino in a load-bearing wall was placed backwards — a sign error. You don't just pull that domino; you walk through every room above it and check whether the lean still holds. That's what OpenAI's math team just did across its entire research corpus. The committed claim here is not a new theorem — it's a new institutional practice. OpenAI's math group discovered that a sign error in "Algebraicity of Weil classes on split abelian eightfolds" invalidated a stabilization-trace cancellation argument. That single bug propagated upward into two dependent papers: one on Kuga–Satake correspondences for K3 surfaces and one on the rational Hodge conjecture for products of K3 surfaces. All three were withdrawn, with public notices linking to archived manuscripts. This is not how most mathematics research groups handle retractions. The scope of the audit is what matters. Beyond the three withdrawals, 14 manuscripts received proof repairs: four papers on Lipschitz heights and Ashkin–Teller currents got crossing, boundary-attachment, conditioning, and convergence argument fixes. Six Kähler minimal model program papers needed expanded positivity and contraction arguments with clarified dependency chains. Two taming/hypersymplectic deformation papers corrected a cone-equality claim and removed an unnecessary dependency. One paper revised torus-projection estimates. One removed an obsolete citation. Then 13 more manuscripts were updated just to point at the corrected versions of companion papers. That's 30 manuscripts touched in a single audit pass. The formalization numbers tell the real story about where this project is headed. OpenAI reports 300 of 719 top-line results now formalized — roughly 42%. Formalization means machine-checked proof, not just peer review. Six new formalizations and five supporting additions came in this update alone. The implicit claim is that formalization is the integrity floor: if you can formalize a result, the sign-error class of bug gets caught mechanically rather than by human inspection months later. The ladder question — how does this compare to standard practice? — is stark. In traditional mathematics, retractions are rare, painful, and often incomplete. Dependent papers frequently remain in the literature with no notice. The practice of publicly mapping dependency chains between your own papers, withdrawing the full cascade, and revising everything downstream is closer to software engineering's dependency-management discipline than to standard academic norms. There is no real baseline to compare against because almost no research group operates this way. The integrity regime here is unusual: the validation is partly self-grading (the authors found their own error and audited their own corpus), but the formalization layer adds an external check — the proof assistant doesn't care who wrote the theorem. The 42% formalization rate means 58% of top-line results still rely on human-only verification, which is where the next sign error could hide. The honest gap is scale. OpenAI has not published the full dependency graph of its 719 results. We don't know how many results depend on the 58% that aren't yet formalized, or what the expected error rate is in those unformalized results. The obvious next experiment — running the same kind of audit on the unformalized 58% with adversarial red-teaming, not just author self-review — wasn't done. Most likely reason: the formalization pipeline is compute-and-labor-constrained, and they're prioritizing the most load-bearing results first.