Imagine you hire a detective to explain how a suspect committed a crime. The detective meticulously reconstructs every law-abiding day the suspect ever had — the commute, the groceries, the gym — and delivers a 300-page dossier that explains 99% of their life. But it accounts for almost none of the criminal acts. You'd fire that detective. This paper argues that mechanistic interpretability's circuit-based explanations are doing exactly this: reconstructing the well-behaved outputs while quietly dropping most of the model's mistakes on the floor. The committed claim is straightforward: circuits validated by standard ablation methods can achieve near-perfect agreement with the full model on correct outputs while failing to reproduce the majority of the model's errors — and this gap undermines the explanatory power those circuits are supposed to deliver. The authors test this across IOI (indirect object identification), Docstring, and six model-task settings from the Mechanistic Interpretability Benchmark (MIB). The numbers are damning in their asymmetry: on IOI for GPT-2 small under mean ablation, manual and automated circuits agree with the model on 97.3–99.5% of correct answers but only 11.4–41.7% of its errors. The architectural family here is ablation-based circuit discovery — the standard MI pipeline where you identify a subnetwork (circuit) within a transformer, ablate (zero, mean, or resample) everything outside it, and check whether the circuit's outputs match the full model's. The key insight is that existing validation metrics (KL divergence on the full distribution, or overall accuracy match) are dominated by the model's successes, which vastly outnumber its failures. A circuit can score brilliantly on aggregate while being almost blind to the specific computations that produce wrong answers. The paper doesn't propose a new discovery algorithm — it proposes a new evaluation requirement: exact error agreement, measured separately on the subset of inputs where the model fails. The ladder comparison is internal rather than against an external SOTA. The authors benchmark manual circuits (Wang et al., 2022), automated circuits from ACDC, EAP, SP, and a distributional circuit trained to match the full model's output distribution (not just its top-1 prediction). Even the distributional circuit — designed to capture the entire output profile — only recovers 41.7% of IOI errors under mean ablation. The strongest result comes from the case study: restoring a small set of omitted attention heads raises error reproduction from 14.2% to 75.1% on a held-out set, with only a 0.41 percentage-point drop in correct agreement. This is the paper's sharpest empirical contribution — showing that error-critical computation lives in specific, identifiable heads that current circuit extraction systematically leaves out. Integrity is solid but bounded. The evaluation uses established community tasks (IOI, Docstring, MIB), with clear splits between correct-answer and error prompts. The control experiments are well-designed: the restored heads outperform matched random extensions and scalar-biased controls, which rules out the obvious confound that any additional capacity trivially recovers errors. However, the entire analysis is on GPT-2 small — a model small enough that MI community norms accept it, but far from the frontier models where explanations would matter most. No code release is mentioned. No pre-registration. The milestone this work points toward isn't a qubit count or an accuracy number — it's a methodological standard. The next concrete target would be: an automated circuit discovery method that achieves ≥90% error agreement AND ≥95% correct agreement simultaneously on IOI, without manual head restoration. No existing method hits this bar. The gap is unclear — this could be a year away if the community adopts error agreement as a metric, or indefinitely far if the field doesn't change its evaluation norms. The obvious experiment not run: scaling this analysis to larger models (GPT-2 medium/large, or a Llama-scale model). The authors don't attempt it, and the honest read is (a) — compute and complexity. Enumerating attention heads and running combinatorial ablations on larger models is expensive, and the MI community's standard testbed is GPT-2 small. But the paper's central argument becomes far more consequential at scale, where we actually need explanations. This is the gap that limits the work's impact.