Imagine you're a bartender working a busy counter. Two people are mid-argument about rent prices, a third is trying to order a drink, and a fourth just switched from English to Mandarin to take a phone call. You track all four threads simultaneously — who's talking to whom, what was said ten minutes ago, when to interject. Now imagine building a single AI that does this. That's the core problem MultiTalk attacks. The committed claim: a single end-to-end speech model can sustain coherent, full-duplex, multi-party, bilingual (English-Chinese) conversation over extended durations — averaging 32.6 minutes — and substantially outperform existing open-source baselines on a new benchmark built from real human recordings. This isn't a methods-only paper. They ship data, benchmark, and a trained model. The field context matters. Full-duplex speech models — where the system can listen and talk simultaneously, handling interruptions and backchannels like a human — have been advancing rapidly since Moshi (Kyutai, 2024). But the entire paradigm has been stuck in a two-person, short-conversation box. Real meetings, classrooms, and social settings involve N speakers over long durations with complex turn-taking, addressee shifts, and long-range coreference ("the thing Sarah mentioned twenty minutes ago"). Nobody had the data to train for this or the benchmarks to measure it. MultiTalk's architecture extends the Moshi paradigm — a codec-language-model that jointly models multiple audio streams at the frame level — along two axes simultaneously: participant count and conversation length. The training pipeline is synthetic: they use LLMs to generate multi-party dialogue scripts with controllable properties (length, speaker count, overlap patterns, code-switching, long-range entity references), then render these to audio via TTS. The 57.6k hours of synthetic data (MultiTalkPT for pretraining, MultiTalkFT for fine-tuning) are released on HuggingFace. The evaluation benchmark, MultiTalkBench, is built from real human recordings — not synthetic — and includes probes for long-range entity tracking, topic coherence, and addressee selection. The ladder is honestly presented. They compare against three named open-source baselines: the original Moshi, MiniCPM-o-4.5, and Qwen3-Omni-30B-A3B-Instruct. MultiTalk "substantially outperforms" all three on MultiTalkBench. The honest caveat: all baselines were designed for dyadic (two-party) conversation and short contexts, so the comparison is somewhat asymmetric — it's less "we beat SOTA" and more "we built the first system that even attempts this task and it works." No proprietary systems like GPT-4o Advanced Voice are benchmarked, which is a notable gap. The integrity picture is mixed in instructive ways. The benchmark uses real human recordings (strong), the training data and benchmark are publicly released (strong), and they compare against multiple named baselines (good). But the training data is entirely synthetic, which means the model may be learning the statistical regularities of LLM-generated dialogue rather than real human multi-party conversation patterns. The benchmark evaluates on real recordings, which partially addresses this concern, but the gap between synthetic training distribution and real evaluation distribution is a known fragility point. No pre-registration, and no independent replication yet — expected for a NeurIPS 2026 submission. The milestone that matters: current conversations average 32.6 minutes with controllable participant count. The next meaningful threshold is probably 60+ minutes with 6+ speakers on real (not synthetic) training data, which would bring this into genuine meeting-transcription-and-participation territory. The experiment they didn't run — and this is the obvious one — is training on real multi-party conversation data rather than synthetic. The most likely reason: real multi-party full-duplex data with proper speaker diarization and codec-frame-level alignment simply doesn't exist at scale. Building the synthetic pipeline was the path of least resistance, and they're probably hoping the community builds the real-data version next.