Imagine a student who always circles the right answer on a multiple-choice exam but whose scratch work is nonsense — invented arithmetic, skipped steps, numbers that don't connect. You'd never trust them to solve a new problem. That's the state of LLM-based vulnerability analysis today, and this paper is the first structured attempt to grade the scratch work. The core claim is uncomfortable: approximately 60% of correct vulnerability verdicts from LLMs are accompanied by fabricated or unverifiable reasoning claims. Chain-of-Thought prompting — the standard practice for getting models to "explain" their security judgments — produces plausible prose that hides logical leaps, hallucinated execution traces, and internal contradictions. Worse, when you use another LLM as a judge (the increasingly popular LLM-as-judge paradigm), the judge systematically misses these errors because it evaluates surface coherence, not logical validity. VERA's intervention is architectural rather than just evaluative. Instead of accepting free-form text explanations, it forces the model to output a Structured Reasoning Record (SRR) — machine-readable fields encoding tracked pointers, memory operations, and state transitions. Think of it as replacing the essay question with a fill-in-the-blanks worksheet where every claimed pointer dereference or buffer state must be explicitly declared. A multi-stage judge then audits each SRR against eight codified reasoning failure modes using deterministic checks first, with LLM calls reserved only for semantic interpretation where rules can't reach. The ladder result matters: VERA exposes 87% of reasoning errors that free-form LLM-as-judge evaluation systematically misses. The paper also surfaces a deeply unsettling finding — reasoning flaws occur in correct verdicts just as frequently as in incorrect ones. This means accuracy metrics alone are fundamentally misleading for security tooling. A tool that's 90% accurate on binary vulnerability classification could still be producing dangerous garbage explanations that mislead the humans doing triage and patch engineering. The validation design includes an automated mutation testing approach that benchmarks judges at scale without requiring human annotation. This is a pragmatic choice — human annotation of vulnerability reasoning is expensive and doesn't scale — but it also means the ground truth for "correct reasoning" is itself synthetic. The paper acknowledges this by defining eight specific failure modes (hallucinated execution steps, fabricated pointer states, etc.) that can be deterministically checked against the SRR schema, reducing circularity. The practical consequence is a direct challenge to the emerging "LLM-as-judge" paradigm in security. If your automated judge can't distinguish correct reasoning from plausible-sounding hallucination, you don't have an automated auditor — you have an automated rubber stamp. VERA's structured schema approach is the kind of constraint that actually makes AI systems more trustworthy by reducing the surface area for plausible-sounding fabrication. What's missing is scale testing across diverse LLM architectures and vulnerability types, and real-world deployment data showing whether VERA's structured format degrades the model's binary classification accuracy (forcing structured output sometimes hurts performance). The authors are likely saving the multi-model comparison for a follow-up, but the absence is notable.