Imagine you're a drummer and a bassist who've never met, thrown on stage for a gig with one shared setlist. The drummer controls tempo and dynamics; the bassist controls harmony and groove. They need to coordinate in real time, but they can't just play the same instrument — each has its own physics, its own action space, its own timing constraints. Most bands solve this by having one player lead and the other follow. But what if you could train both musicians simultaneously on thousands of recordings from different bands, different genres, different instruments, so that when they finally take the stage together, they share a deep intuitive sense of what the other is about to do — even though they never rehearse the specific song? That's the core mechanism of MM-ABC. The paper's committed claim: a single vision-language foundation model can coordinate a mobile robot's base locomotion and arm manipulation through separate but cross-attending action streams, pretrained on 5,000+ hours of heterogeneous data spanning 17 embodiments. This isn't the first paper to tackle mobile manipulation, but it's a serious bid for the first generalist architecture that treats arm-base coordination as a first-class design problem rather than an afterthought. The architecture has three interlocking pieces worth understanding. First, 'Seeing' — sparse multi-level features extracted from a vision-language model, giving the system spatial grounding that survives ego-motion (the camera moves because the robot moves). Second, 'Imagining' — a training-only future branch that forces the model to predict what the world will look like next, using world imagination and geometric intent as auxiliary supervision signals. This branch is discarded at inference time but leaves behind a representation that better understands manipulation intent. Third, 'Coordinating' via MM-APT — separate manipulation and mobility streams linked through masked joint attention and a clean-action cross-prediction mechanism. The clean-action prediction is load-bearing: replacing it with standard velocity prediction drops success on composite tasks from 32.8% to 29.2%. The ladder comparison is where things get interesting and where you should calibrate your expectations. MM-ABC hits 44.71% on EBench (a mobile manipulation benchmark), 61.2% on RoboCasa365, 99.1% on LIBERO, 82.8% on LIBERO-Plus, and 83% on five real-world tasks. The LIBERO numbers are strong — near-ceiling. But EBench at 44.71% and RoboCasa365 at 61.2% tell you we're still far from reliability for deployment. The paper's ablations are its strongest integrity feature: they systematically remove each component (future supervision, multilevel conditioning, clean-action prediction) and show measurable drops, which is exactly the evidence you want. The pretraining scale is genuinely impressive — 400K+ episodes across 12 datasets — making this one of the larger robot foundation model efforts published. Integrity-wise, the benchmarks are community-standard (EBench, RoboCasa365, ManiSkill-HAB, LIBERO), which is good. The real-world evaluation covers five tasks at 83% mean success, though the number of real-world trials isn't specified in the abstract. The ablation design is clean: controlled single-variable removals on composite-seen tasks. No pre-registration, no independent replication yet, and the code/model availability status is unclear beyond the project webpage. The milestone that matters for this line of work is crossing from ~60% success on compositional household tasks to ~90%, which is the rough threshold where a robot becomes useful rather than entertaining. MM-ABC is at 61.2% on RoboCasa365. If the scaling curves from pretraining data hold, doubling or tripling the data and adding more embodiments could close the gap, but we've seen enough robotics scaling curves plateau to know that's not guaranteed. Watch for whether 10K+ hours of pretraining data yields proportional gains. The obvious experiment not run: zero-shot transfer to a completely novel embodiment not in the pretraining mix. With 17 embodiments in training, the generalization claim is implicit but untested. My honest read: this is being saved for the next paper, because it's the most publishable result if it works and the most damaging omission to report if it doesn't.