You know how a good driving instructor doesn't need to know the "correct" route to evaluate whether you can drive? They watch for coherence: do you check mirrors, signal before turning, respond to speed limits proportionally? They're testing internal consistency against known rules of competent driving, not scoring you against a single right path. That's exactly the mechanism this paper ports from stated-preference economics into LLM evaluation. The committed claim: the validity framework that environmental and health economists have used for decades to evaluate survey responses — content validity, construct validity, criterion validity, reliability, incentive compatibility, consequentiality — is a general-purpose evaluation method for language models on questions that have no correct answer. The authors aren't proposing a new benchmark. They're proposing a new category of evaluation logic. To demonstrate, they administer a published water-quality valuation survey (Vossler et al. 2023) to six LLMs across model generations. The tests are predictions from economic theory: demand should slope downward (people should want less of something as it costs more), willingness to pay should increase with the scope of the good being valued, and it should respond to income. These aren't arbitrary — they're structural consistency checks derived from decades of microeconomic theory. The results separate models with surgical clarity. Two older models fail the most basic demand-slope test at $75,000 household income. The two newest models pass every theoretical validity test the authors can score, but diverge on convergent validity — meaning they're internally coherent but don't agree with each other on the actual numbers. The intellectual architecture here is framework transfer, not algorithm design. The paper belongs to the family of evaluation methodology papers, not model-building papers. It leans on the accumulated institutional knowledge of stated-preference economics — contingent valuation, discrete choice experiments, the entire post-NOAA-panel literature on survey design — and argues this infrastructure is portable. The key structural insight is that "validity" in this tradition is not binary pass/fail but a multi-dimensional profile: a model can have strong construct validity (its answers obey the right theoretical relationships) while having weak criterion validity (its answers don't match what real humans say). The integrity picture is mixed but honest. The authors use a real, published survey instrument rather than inventing their own, which removes one layer of circularity. They test six models, enough to show separation but not enough to draw robust generalization curves. The theoretical predictions they test against (downward-sloping demand, scope sensitivity, income effects) are well-established in economics but represent only one domain's consistency checks. The paper is transparent that passing validity tests shows coherence, not correctness — a distinction most LLM evaluation papers gloss over or ignore entirely. The real contribution is conceptual, not empirical. The empirical demonstration is a proof of concept for a much larger argument: that any domain with well-developed theoretical predictions can use those predictions as validity tests for LLM outputs, even when no ground truth exists. This reframes LLM evaluation from "how close to the right answer" to "how structurally coherent is the response profile" — a move that matters enormously for policy, ethics, and preference elicitation applications where right answers don't exist. What's missing is the scaling question. The authors demonstrate the framework on one survey in one domain (environmental economics). The obvious next experiment is applying the same validity taxonomy across multiple domains — health economics, political preference elicitation, moral reasoning — to see whether the framework's discriminating power holds or whether it's domain-specific. The authors likely didn't run this because it would require deep collaboration with domain experts in each field, not because the experiment would fail. This is a "saving it for the research program" situation, not a hidden negative result.