Imagine you're playing Wordle, and after one letter turns green you shout the answer — and you're right. You feel like a genius, but your process was garbage. Now imagine a doctor doing the same thing: pattern-matching a diagnosis from a single lab result before the CT scan, the biopsy, or the second opinion arrives. The answer might be correct, but the reasoning path is a clinical landmine. This paper names that failure mode Evidence-Value Misalignment (EVM) and builds a benchmark to measure it. The core claim is sharp: outcome-based accuracy — the standard way we evaluate medical LLMs — actively rewards lucky guesses and hides dangerous reasoning failures. MedEVM, the benchmark, presents 1,050 cases across 24 disease systems as sequential turn-by-turn observation streams. At each turn the model must decide: do I have enough evidence to commit, or should I wait? This is the diagnostic muscle that matters in real clinical settings, and it's precisely what accuracy-only benchmarks cannot test. Four findings land hard. First, models are badly miscalibrated about evidence sufficiency — they cannot tell when they have enough to commit, and reasoning-mode models (chain-of-thought, etc.) actually get worse at this, not better. Second, even when a model is confident in the right answer AND has sufficient evidence, it frequently fails to submit the diagnosis — a bizarre hesitation bug. Third, reordering identical evidence changes the diagnosis even when model confidence stays flat, meaning the models are sequence-sensitive in ways clinicians should not be. Fourth, misleading evidence injected after sufficient evidence has already arrived still redirects diagnoses — the models cannot maintain conviction when challenged by noise. The validation design deserves credit. Nine LLMs tested. The benchmark forces dynamic interaction rather than static QA. The authors explicitly measure the gap between 'correct diagnosis' and 'correct diagnosis supported by sufficient evidence,' which is a genuinely novel decomposition. They also show that EVM scores predict actual diagnostic errors, giving the metric real teeth. The proposed fix, EVD-Harness, decouples diagnosis generation from submission through three online control stages plus an offline Contrastive Diagnostic Wiki. Across five LLMs, this harness improves accuracy by 12.0 to 51.1 percentage points while specifically reducing EVM-related failures. The architecture is a pipeline wrapper, not a new model — it works by preventing premature commitment and verifying evidential support before allowing submission. Think of it as a mandatory checklist inserted between the doctor's hunch and the signed chart. The integrity picture is mixed. The benchmark is custom-built by the authors, not a community standard yet. No code availability is mentioned. The 9 LLMs tested are not all named in the abstract (the full paper presumably names them). The comparison is models-with-harness vs models-without, not against a clinical gold standard or against existing medical AI benchmarks like MedQA or USMLE. The 12–51 point improvement range is wide enough to suggest high variance across models. What makes this paper matter beyond its specific results is the conceptual wedge it drives between accuracy and evidential grounding. If you accept EVM as a real failure mode — and the four behavioral findings are convincing — then every medical LLM benchmark that reports only accuracy is systematically overstating clinical reliability. That's a claim with legs.