Imagine two chefs sharing a kitchen who don't speak the same language. One solution: they each shout raw sensor data — exact arm positions, gripper forces, pixel arrays — across the counter. That's bandwidth-heavy and fragile. The smarter solution: they learn to say 'I'm plating the fish, you start the sauce.' That's semantic communication — compressed, contextual, actionable. DuoMind is the second approach, applied to robots. The core claim: a distributed hierarchical framework where each robot runs two stacked foundation models — a vision-language model (VLM) as the 'orchestrator' doing high-level reasoning and inter-agent messaging, and a vision-language-action model (VLA) as the 'executor' turning those instructions into motor commands. At each planning step, the orchestrator ingests the task instruction, its own camera feed, and semantic messages from its partner, then emits both a low-level action instruction and a natural-language message to the other robot. No centralized planner. No shared state server. Each robot is a self-contained reasoning-and-acting unit that coordinates through language. The architectural bet is clean: VLMs are good at reasoning and communication but bad at precise action generation; VLAs are good at fine-grained manipulation but lack multi-step planning and coordination capacity. DuoMind stacks them so each does what it's best at. The semantic messages are the key differentiator — rather than broadcasting full observations or low-level plans, robots exchange compressed natural-language descriptions of intent, state, and requests. This is bandwidth-efficient and, in principle, more robust to observation noise than raw data sharing. To test this, the authors built RoboPoly, a new benchmark of long-horizon bimanual manipulation tasks requiring coordinated, closed-loop execution under distributed control. They also evaluate on the existing RoboTwin benchmark. DuoMind outperforms baselines on both, and ablation studies show that both the hierarchical split (VLM over VLA) and the semantic communication channel contribute independently to performance. The ablations are the strongest part of the empirical section — they isolate the communication and hierarchy effects rather than just reporting aggregate numbers. The integrity picture is mixed. RoboPoly is the authors' own benchmark — necessary because multi-robot coordination benchmarks are scarce, but also means they're grading their own exam. The RoboTwin evaluation provides a second data point, but both environments are simulated. No real-robot experiments appear. The baselines compared are reasonable (single-robot VLAs, centralized planners, communication-free multi-robot setups), but the field lacks a clear SOTA multi-robot coordination system to ladder against. The paper is honest about being simulation-only, but this is the load-bearing gap. The milestone question is whether semantic communication between foundation-model agents can survive the sim-to-real gap — noise, latency, partial observability, and the messiness of physical manipulation. The paper demonstrates the architecture works in simulation with two robots; the next number to watch is whether it transfers to physical dual-arm setups with even moderate task complexity (5+ step horizons, real objects, real sensors). That's likely 1-2 years away if the team has hardware access. The obvious experiment not run: real-robot deployment. The honest read is (a) — they likely don't have the hardware setup or the engineering budget to close the sim-to-real loop for this submission. Scaling beyond two robots is the second gap; the framework is described as distributed but only tested with pairs. Three or more agents introduce coordination complexity (message routing, conflict resolution, partial observability cascades) that pairwise messaging doesn't capture.