Imagine you want to rank chefs by skill, so you give nine different food critics the same fifteen chefs to evaluate. But one critic only rates desserts, another only rates sushi, and a third penalizes anyone who refuses to cook foie gras. You'd expect the rankings to diverge — not because the critics are bad, but because they're measuring different things under the same label. That's exactly what happens when you try to benchmark LLM persuasion. The committed claim here: nine published automated persuasion evaluation methods, adapted to a shared setup and run on the same fifteen LLMs, agree only weakly — mean Spearman ρ = 0.25. That's barely above noise. The paper doesn't just report the disagreement; it decomposes it. Two factors dominate. First, model refusals: LLMs that decline manipulation tasks but comply with rational persuasion tasks scramble the rankings, accounting for roughly a quarter of the disagreement. Second, general capability: non-manipulative persuasion methods largely track general LLM capability (smarter models persuade better), while manipulation methods do not. These are different constructs masquerading under one label. The architecture is methodological rather than algorithmic. The authors take nine existing evaluation protocols — spanning prompt-based opinion shifts, debate formats, and manipulation scenarios — normalize them to a shared model set, and compute rank correlations across all pairwise method combinations. The compute is modest: API calls to fifteen LLMs across nine evaluation pipelines. The analytical backbone is rank-order statistics, not deep learning. This is a meta-evaluation, not a new model. The integrity picture is mixed but honest. The authors openly acknowledge that with nine methods (reduced to eight for some analyses after exclusions), statistical power for explaining disagreement is limited. They call their decomposition 'indicative' rather than definitive. The fifteen models span a reasonable range of capability but are not named in the abstract — the reader needs the full paper to assess selection bias. No pre-registration, no code availability mentioned. The validation is inherently same-team: the authors chose which methods to include and how to adapt them. The practical upshot is sharp. Persuasion scores are not portable across evaluation contexts. A model that scores high on rational argumentation may score low on manipulation — not because it can't manipulate, but because it won't. This means regulators or developers relying on a single persuasion benchmark are measuring a joint distribution of ability and alignment, with no way to separate the two from one number. The paper frames this as a measurement problem, not a safety problem, but the safety implications are obvious: a model could ace one benchmark while being dangerous on tasks that benchmark doesn't cover. The field fight this paper enters is whether persuasion is a unitary capability or a context-dependent composite. The dominant assumption in AI safety evaluation has been that you can meaningfully rank models on 'persuasiveness' as a trait. This paper argues that assumption is wrong — or at least unsupported by current methods. It's a measurement critique, not a capability result, and those are often more important for how a field develops. What's missing is the obvious follow-up: factor analysis or latent-variable modeling that explicitly decomposes persuasion scores into ability, willingness, and task-type components. The authors gesture at this decomposition qualitatively but don't build a formal model. The honest read is (a) — this is a workshop-scale effort that ran out of scope, not a hidden negative result. The paper sets up the problem cleanly; the solution paper is next.