Imagine two jazz musicians who can hear each other perfectly but can only start playing after the other one stops. They'd still lock into a rhythm — their pauses would correlate, their entries would be coordinated — but the music would feel sluggish because neither player is anticipating the other's phrasing. That is what this paper finds when two full-duplex speech models have an unscripted conversation. The core claim: when two PersonaPlex-7B instances talk to each other on a shared clock, their turn-taking timing is genuinely coupled — shuffle who talks to whom across conversations and the coupling breaks. But the floor changes hands at a median of 400–560 ms, versus 137 ms for human speakers in Switchboard. That gap isn't noise. It's structural. The last 120 ms of a partner's turn, where humans place roughly 10% of their floor transfers through predictive projection, holds just 1% of the model's transfers. The models are waiting for confirmed silence, not anticipating the end of a turn. The experimental design is clean and minimal. One floor-transfer rule is applied identically to the model dyads and to human Switchboard data, making the comparison apples-to-apples on that metric. The key manipulation is a channel-delay experiment: artificially delaying one direction of the audio stream shifts the model's response onset one-for-one and leaves the pre-offset window empty. In humans, projection-based turn-taking would partially absorb that delay. The models absorb none of it — they simply wait longer. This matters beyond curiosity because the use cases for model-to-model speech are multiplying fast: self-play data generation, agent societies, model-based evaluation. If the models' turn-taking timing is fundamentally reactive rather than projective, every downstream artifact inherits that latency signature. Synthetic dialogue data generated this way won't contain the sub-200ms transitions that characterize natural conversation. Evaluators built on model-to-model interaction will never test a model's projection ability because neither side exercises it. The baseline comparison is honest but limited. Switchboard is the only human reference, and 137 ms is cited as the human median without drilling into variance across Switchboard's demographic and recording conditions. The authors don't compare against other full-duplex models (dGSLM, SpiritLM, Moshi) which would clarify whether the 400–560 ms gap is a PersonaPlex property or a general architectural one. That's the most conspicuous absent experiment. The paper is a diagnostic, not a solution. It identifies a specific failure mode — reactive timing replacing projective timing — and provides a falsifiable signature (the empty last-120ms window, the one-for-one delay shift). What it does not offer is a mechanism to close the gap or a clear argument for why current architectures can't learn projection. The implicit hypothesis is that autoregressive next-token prediction over audio tokens doesn't incentivize turn-end projection the way human conversational pressure does, but this remains speculation. For anyone building or evaluating speech agents, the practical takeaway is concrete: if your model-to-model pipeline generates training data or evaluation scores, check whether its floor-transfer distribution matches human distributions. If median gap-to-next-turn exceeds ~200 ms and the last 120 ms window is empty, your pipeline is producing reactive dialogue, not naturalistic dialogue. That's a measurable, actionable bias.