Imagine judging a tennis player by having them hit balls against a wall. You'd learn their stroke mechanics, maybe their power, but nothing about their ability to read an opponent, adjust mid-rally, or exploit a weak backhand. That's essentially how full-duplex spoken dialogue models — the kind that listen and talk simultaneously, like a real human — are evaluated today. The examiner is either a pre-recorded clip that can't react, or a bot that runs a fixed script and never gets graded itself. This paper argues that's evaluating half a conversation and calling it the whole thing. The committed claim: turn-taking, interruption, and overlap are emergent properties of TWO coupled speakers, not properties of one speaker measured against a static prompt. Therefore, you need dyadic evaluation — two models talking to each other under assigned roles, both scored by an external judge. The authors call this DyaFDB (Dyadic Full-Duplex Benchmark) and instantiate it with four task types across 140 scenarios, producing 7,560 recorded conversations from six self-play and cross-play pairings of current full-duplex models. The architecture is straightforward by design: this is a benchmark paper, not a model paper. Two full-duplex models are placed in conversation with assigned cooperative or conflicting goals (think: one model is a customer service agent, the other is a frustrated caller). An external LLM judge scores both sides offline on dimensions like turn-taking behavior, role adherence, and goal completion. The key structural choice is that no pre-recorded audio is used — every conversation is live, meaning each model's behavior is shaped by the other's responses in real time. The most interesting empirical finding is also the most intuitive once you hear it: how a model behaves continually reshapes its partner. This isn't just a theoretical claim — they observe it across the 7,560 conversations. A model that interrupts aggressively elicits different behavior from its interlocutor than one that waits patiently. This means single-sided evaluation doesn't just miss information; it measures a fundamentally different thing than what happens in real deployment, where the user adapts to the agent and vice versa. The ladder question is tricky because DyaFDB isn't competing against other models — it's competing against other evaluation frameworks. The current standard is single-sided: projects like the Spoken-LLM benchmarks, AudioBench, or custom examiner-bots. DyaFDB doesn't claim to replace these; it claims they're structurally incomplete. The evidence is observational rather than mathematical: the authors show that model rankings shift when you move from single-sided to dyadic evaluation, meaning the two frameworks disagree about which model is better. That's the core argument for why dyadic evaluation is necessary, not just nice-to-have. Integrity is mixed. The scenarios and protocols will be released, which is good. The external judge is an LLM, which introduces its own biases — but the authors acknowledge this. The six model pairings cover self-play and cross-play, which is the right design. However, the benchmark is self-validated: the authors designed the scenarios, ran the conversations, and chose the scoring criteria. No independent group has replicated the framework or confirmed that the ranking shifts they observe are robust to different judge models or different scenario sets. The real milestone here isn't a number — it's adoption. If the spoken-dialogue community shifts from single-sided to dyadic evaluation as a standard practice, that changes what gets optimized. Models would be trained not just for fluency in isolation but for adaptive conversational competence — reading the other speaker, managing overlap, negotiating turns. The successor experiment the authors didn't run is the obvious one: training a model specifically to optimize for dyadic metrics and showing it outperforms models trained on single-sided benchmarks when deployed with real humans. That's the experiment that would close the loop. My read: they're saving it for the next paper, because it requires a training run, not just an evaluation framework.