Imagine you're judging a baking contest, but every judge uses a different thermometer, a different definition of "done," and one judge's oven was set to Celsius while another's was in Fahrenheit. You'd get wildly different winners — not because the cakes are different, but because the measurement rigs are. That's the state of unsupervised anomaly detection in brain MRI, and MIRTO is the paper that finally puts all the thermometers on the table. The committed claim: there is no single honest score for unsupervised anomaly detection (UAD) in brain MRI because the score you get depends on at least four upstream choices that almost nobody reports — how anomaly maps are spatially aligned to the reference brain, where and how the detection threshold is set, which metric and aggregation are used, and what counts as a "lesion" in the first place. MIRTO makes all of these explicit, runs the full combinatorial space (15,552 pipelines), and quantifies how much each choice matters. This is a metrology paper, not a methods paper. The headline finding is devastating for anyone who has ever trusted a single leaderboard number. An axis-order mismatch in registration dropped a diffusion model's voxel AUROC from 0.873 to 0.583 — a 0.29 swing from what amounts to a data-handling bug. Slice-level AUROC barely moved, meaning you'd never catch it unless you checked at the right resolution. The method identity explained ≥0.95 of variance in voxel AUROC and AUPRC, but only 0.14 in lesion sensitivity, where the lesion definition and hit criterion dominated. Translation: when measuring "did you find the lesion," what you call a lesion matters more than which algorithm you used. MIRTO's architecture is not a neural network — it's a structured evaluation protocol. It gates every comparison with a registration-quality check and label-free diagnostics of known statistical power. Thresholds are set on validation data only. False-positive volume realized on test is reported explicitly. Every comparison runs across the full multiverse of defensible pipelines. Statistical inference uses paired subject-bootstrap intervals with multiplicity correction. The machinery here is statistical and computational, not learned. The integrity story is refreshingly honest but also self-aware of its own limits. The authors tested nine hypotheses against explicit criteria, but because the same BraTS 2020 cohort (312 subjects) served to develop the protocol, they label ALL inference exploratory. This is the right call and it is rare. They found that a Dice advantage significant at validation thresholds vanished when false-positive burden was equalized — and they traced this exactly to threshold transfer, not to the method itself. They also showed a training-free modification to REFLECT's latent aggregation that raised Dice by 0.052 at equal false-positive burden, demonstrating the protocol can find real improvements, not just tear down existing claims. The obvious next experiment is applying MIRTO to a larger zoo of UAD methods — the paper tests only four — and especially to methods that have published claims on BraTS or other public benchmarks without reporting registration details or threshold provenance. The authors almost certainly scoped down to four methods for tractability: 15,552 pipelines × more methods × bootstrap resampling is compute-expensive. But the real test of MIRTO's utility is whether it changes published rankings when applied retroactively. That experiment would be uncomfortable for some groups, which may also explain its absence. This paper matters because the field of medical-image anomaly detection is growing fast and the benchmarking infrastructure has not kept up. If MIRTO or something like it becomes standard practice, published AUROC numbers will get harder to game and registration artifacts will get caught before they become paper-mill fuel. If it doesn't get adopted, the field continues flying blind on a measurement stack that can swing results by 0.29 AUROC from a single bug. The protocol itself is fully specified and reproducible — no learned weights, no proprietary data — which gives it a real shot at becoming infrastructure.