Imagine you're testing a chess engine, but the person moving the opponent's pieces keeps accidentally revealing their strategy three moves early. The engine looks brilliant — it wins more games, uses fewer moves — but you're not actually measuring the engine anymore. You're measuring how much the sparring partner leaked. That is the core problem this paper diagnoses in LLM agent benchmarks. UserProxyBench introduces a straightforward but overdue idea: score the user simulator, not just the agent. Built as an evaluation layer on top of the tau-bench family of enterprise tasks, it proposes the User Fidelity Score (UFS), a rubric-based metric that measures whether the simulated user followed its private instructions — independent of whether the agent eventually completed the task. The authors hold the agent constant (GPT-5.5) and vary only the user proxy across 375 enterprise tasks. The result: mean task reward swings by 15.2 points depending solely on which LLM plays the user. That is an enormous confound for any benchmark that treats the user as interchangeable plumbing. The dominant failure mode they identify is premature disclosure: the simulated user volunteers information before the agent asks for it. This is the chess-opponent-leaking-strategy problem. It doesn't hurt task reward — in fact, among successful episodes, it causes the agent to make 1.06 fewer tool calls on average. The agent still 'wins,' but the interaction that produced the win is structurally different from the one the benchmark intended to measure. You're scoring a shortcut, not the capability you designed the benchmark to test. The paper's most useful contribution is empirical: across seven user proxies, they map a cost-fidelity frontier. Cheaper models leak more, expensive models are more faithful, and the relationship is legible enough that practitioners can pick the cheapest simulator that meets a required fidelity threshold. This is practical infrastructure for anyone building multi-turn agent evaluations at scale. Architecturally, UFS is a rubric-based scoring method — task-grounded criteria evaluated by an independent judge, not a learned reward model. This keeps the evaluation interpretable and avoids the circular validation trap where another LLM scores itself. The rubric criteria are tied to the benchmark's private user instructions, which means UFS measures specification adherence, not conversational naturalness or some vague 'helpfulness' proxy. The integrity picture is decent but bounded. The agent is fixed at GPT-5.5, the benchmark family is tau-bench, and the sample is 375 tasks — meaningful but not exhaustive. There's no independent replication, no pre-registration, and the authors don't test whether UFS generalizes to non-enterprise benchmarks or non-tau-bench task structures. The 24.4% violation rate among successful episodes is the most provocative number in the paper, but it comes from a single agent-benchmark pairing. The real value here is conceptual, not methodological. The paper forces the field to confront a measurement hygiene problem: if your multi-turn agent benchmark doesn't score the user, you don't know what your benchmark is measuring. The 15.2-point reward swing is the number that should keep benchmark designers up at night.