Imagine you're a teacher grading multiple-choice exams, and you discover that students who randomly shift every answer by one letter — A→B, B→C, C→D, D→A — score just as well as students who carefully studied the material and changed specific answers. You wouldn't conclude the random-shifters learned something. You'd conclude the test is broken. That's the core finding here. The paper attacks a specific claim in adaptive computation research: that oracle evaluations of layer programs (schemes that skip or repeat transformer layers) demonstrate meaningful headroom for input-dependent inference optimization. When researchers use known answers to pick the best layer program per question, they see gains of 9–10 percentage points over fixed baselines. The field interprets this as evidence that smarter selectors could unlock real performance. Guo and Liu show this interpretation is premature. Their experimental design is elegant. They test 32 layer-skipping and repetition programs across Qwen3-4B-Base and Llama-3.1-8B on 4,413 multiple-choice items. The key move: they construct input-blind controls — random-direction perturbations applied at the same network sites — that have no relationship to the input at all. These controls produce 10.2–11.8 and 15.6–19.4 percentage points of headroom, exceeding the real programs' 9.0 and 10.1 points in every random draw. Something that knows nothing about the input shouldn't outperform something designed to adapt to it. The mechanism is option-order sensitivity. Multiple-choice benchmarks present answers in a fixed ABCD layout. Any perturbation that changes which letter the model favors can look beneficial when you retroactively select the best perturbation per question. Fixed letter offsets — just shifting all answers by a constant — produce headroom of similar scale, confirming this is format leakage. When the authors rotate option order to break this coupling, both real and control headroom collapse sharply. The residual real-minus-control gap after rotation is 1.4–4.5 points depending on model and adjustments, but statistical significance depends on corrections and reference choices. There are important caveats. A supplementary generated-answer test (free-form, no ABCD options) finds that search-selected programs retain a 26-point advantage over programs selected for other problems, even after rewording. But this test lacks a placebo comparison — the random-perturbation control wasn't run — so it's suggestive rather than dispositive. KL-calibrated comparisons, including an input-dependent control, favor real programs in point estimate but with inconclusive corrected tests. The evidence is genuinely mixed once you look past the headline. The real contribution is methodological: oracle evaluation on MCQ benchmarks is not a valid way to estimate adaptive-computation potential. The field needs either format-independent evaluation (generated answers with proper controls) or much more careful isolation of what the oracle is actually selecting on. Every paper reporting oracle headroom on MMLU-style benchmarks should now include input-blind controls as a sanity check. This doesn't kill adaptive computation as a direction. It kills one specific evidence pipeline that the community was treating as load-bearing. The distinction matters: the question of whether input-dependent layer selection can improve inference efficiency remains open, but the evidence base just got thinner.