Imagine you hire an accountant. They produce a beautiful tax return with a confident bottom line — but you have no idea whether the numbers on each line actually came from your receipts. You could hire a second accountant to check the final number, but the smarter move is to build a system that traces every line item back to its source document. That's what this paper does for medical AI: it doesn't try to make better predictions. It builds the receipt-tracing layer. The committed claim: verifiability — whether each atomic statement an AI makes about a patient is grounded in that patient's actual radiomic measurements — can be engineered and evaluated as a completely separate module from predictive accuracy. The authors demonstrate this by building a neuro-semantic verification framework for glioblastoma radiogenomics, converting 1,728 radiomic features from MRI scans into machine-checkable 'evidence records' with deterministic provenance. When they deliberately degrade the prediction model's performance from AUC 0.899 down to coin-flip 0.500, the verification layer maintains 100% accuracy on 6,620 corrupted claims. The verifier doesn't care whether the prediction is right — it cares whether each statement traces to the data. The architecture is a pipeline, not a monolith. Radiomic features extracted via CaPTk from four MRI modalities (T1, T1GD, T2, FLAIR) across three tumor sub-regions get converted into 'semantic states' — discretized bins derived from 611 UPenn-GBM reference cases. These states become an evidence ledger: 1,655 model-linked records for 331 external patients. A frozen deterministic verifier then checks whether any generated claim — from an LLM or otherwise — is consistent with the ledger. There's no neural network in the verification step. It's lookup and logic. Cross-cohort transportability is the stress test, and it's honest about the cracks. Median semantic-state agreement between the UPenn reference and the independent multicenter cohort was 0.786 (weighted kappa 0.709), but the range is enormous: morphologic features transport at 0.918, intensity features at just 0.252. The paper correctly flags intensity features as fragile across scanners and sites — a known problem in radiomics that they don't paper over. The MGMT methylation prediction task (AUC 0.543 externally) is explicitly framed as a transport stress test, not a clinical claim. The LLM pilot is the most forward-looking piece: GPT-5.6 Sol reproduced 72/72 prespecified atomic claims in a 24-case test, and the frozen verifier recovered all 24 expected conditions. It's tiny — 24 cases, prespecified claims — but it demonstrates the division of labor the authors are proposing: LLMs do the natural-language structuring, deterministic logic does the evidence checking. The paper is honest that this is a proof of concept, not a deployment-ready system. What's missing is any engagement with the claim-generation failure modes at scale. Twenty-four cases with prespecified claims is a controlled demo; the interesting question is what happens when an LLM generates novel, unexpected, or subtly wrong claims that aren't in the prespecified set. The paper also doesn't benchmark against any existing explainability or fact-checking frameworks (LIME, SHAP, or the growing literature on LLM grounding). The ladder is thin because the authors are essentially defining a new evaluation category rather than competing in an existing one. The real contribution here is conceptual, not empirical: the separation principle. If the field adopts the idea that verification should be evaluated independently of prediction — that a model can be wrong but its explanations should still be checkable — that changes how we think about deploying AI in clinical settings. It's a small paper with a big architectural idea, and the integrity of the demonstration is higher than most preprints in this space.