Imagine you're a driving instructor. You have two students who both run a stop sign. Instructor A says, 'That was wrong. Stop at stop signs.' Instructor B says, 'I see why you thought the road was clear — and I want you to get home safe, which is exactly why stopping matters here.' Both deliver the same verdict: stop. But Instructor B is more likely to be heard. The problem is that if you're grading instructors on a 'pushover scale,' Instructor B looks softer. This paper says that's exactly the bug in how we evaluate AI chatbots. The committed claim: current social-sycophancy benchmarks in LLM evaluation are construct-invalid because the linguistic markers they penalize — hedging, validation, acknowledgment — are the same markers that define conversational receptiveness, a well-established social psychology construct shown to improve persuasion and trust across disagreements. The paper doesn't just argue this theoretically; it demonstrates the overlap empirically on a widely used moral-advice dataset and then runs a preregistered human-subjects experiment to show that users actually prefer receptive responses and find them more persuasive, even when those responses disagree with the user's initial position. The architecture here is methodological rather than algorithmic. The authors take an existing social-sycophancy evaluation framework (applied to a moral-advice dataset), measure conversational receptiveness on the same responses using validated scales from social psychology, and show tight positive correlation: responses scored as more sycophantic are also scored as more receptive. They then run the critical causal test — manually rewriting human responses to increase receptiveness while holding the substantive conclusion fixed — and show that this intervention alone causes the sycophancy classifier to flag them harder. That's the kill shot for construct validity. The preregistered experiment is the load-bearing wall. Participants (sample size not specified in the abstract, but preregistration is stated) compared substantively equivalent responses that varied only in receptiveness. The more receptive versions won on every metric: preference, expected persuasiveness, and willingness to seek future advice from the author. This held even among participants who believed the original question-asker was morally wrong — meaning the preference for receptiveness isn't just people liking agreement. It's people liking engagement. The paper closes with a practical contribution: a simple method to increase receptiveness without increasing substantive deference. This is the decoupling proof. If you can dial up the social warmth without dialing up the agreement, the two constructs are separable — and evaluation frameworks that conflate them are measuring the wrong thing. The details of this method aren't specified in the abstract, but its existence is the paper's strongest card for applied impact. On the integrity front, the preregistration is a genuine strength — it's relatively rare in NLP/AI evaluation papers and signals the authors anticipated the 'we chose metrics after seeing results' attack. The main limitation visible from the abstract is scope: one dataset (moral-advice), one cultural context, and we don't know the sample size or demographic composition of the human experiment. The field fight this enters — whether LLM alignment evaluations are measuring what they think they're measuring — is live and consequential, because these evaluations directly shape RLHF training signals. The twenty-year version of this question matters enormously. If sycophancy benchmarks are penalizing receptiveness, then RLHF pipelines trained to minimize sycophancy may be producing models that are blunter, less persuasive, and less effective at helping users update their beliefs. The paper doesn't just identify a measurement bug — it identifies a potential alignment tax where fixing one problem (deference) creates another (dismissiveness).