Imagine you're a teacher who gives students a single grade for "math ability" — but the test mixes calculus, geometry, statistics, and arithmetic, and most students ace the arithmetic so thoroughly that the score is really just measuring the other three. Worse, the remaining questions don't test one skill — they test several at once, and students from different schools score differently on the same question even when their overall ability is identical. That's not a math test. That's noise wearing a label. This paper catches HELM Safety doing exactly that. The committed claim: "harmful refusal" — a model's tendency to refuse dangerous prompts — does not function as a single measurable attribute in HELM Safety, and therefore the benchmark's scores cannot be meaningfully compared across models. The authors import construct validity methodology from psychometrics (the science of test design) and apply it to AI safety evaluation, arguing that before you can measure something, you need to demonstrate that the something actually exists as a coherent, unitary construct in your test. The demolition is systematic. Of HELM Safety's four datasets that plausibly target harmful refusal, three are saturated — models score so high that there's no meaningful variance left to measure. Only HarmBench survives this initial filter. The authors then subject HarmBench to two formal tests. A multidimensional item response theory (MIRT) model strongly suggests the dataset is not measuring a single latent trait — it's measuring several collapsed together. A differential item functioning (DIF) analysis reveals items where models from different developers with the same overall refusal ability score differently, suggesting developer-specific biases or training artifacts contaminate the measurement. The DIF flags largely disappear when the authors match models within specific harm scopes rather than across the full dataset — a pattern consistent with Simpson's paradox-style aggregation effects. But this finding cuts both ways: it means the aggregate score is actively misleading, conflating behaviors that differ by harm category. The paper is careful not to overclaim here, noting that scope-specific matching cannot fully rule out genuine developer-level differences in how models handle specific domains. The architectural contribution is methodological, not computational. The authors transplant a well-established measurement science framework — construct validity, IRT modeling, DIF analysis — into an AI safety evaluation context where these tools are almost never used. The compute requirements are modest; the intellectual contribution is showing that the emperor has no clothes. HELM Safety's top-line number collapses HarmBench scores with three other saturated datasets into a single aggregate, meaning the final safety score is dominated by ceiling effects and multi-dimensional noise. The practical consequence is sharp: model comparisons based on HELM Safety's aggregate score are not meaningful. Two models with identical top-line safety scores can have radically different profiles across harm categories, and the score structure provides no way to detect this. The paper argues that any safety score should "earn its single-attribute reading" before being used comparatively — a standard that most current AI safety benchmarks would fail. This matters beyond HELM. The same pathology — saturation plus multidimensionality plus aggregation — likely afflicts other safety benchmarks that collapse diverse harm types into single scores. The paper doesn't audit those benchmarks, but the toolkit it demonstrates is directly portable. The next time someone publishes a leaderboard ranking models by "safety," you should ask: did anyone check whether that number measures one thing?