Imagine you're a wedding planner who has never visited the venue. You have two separate experts: one who knows the venue's floor plan — where chairs fit, which doors open — and another who knows exactly how to choreograph a first dance. Neither expert has worked a wedding at this specific venue before. Your job is to get the dance right in the space by having the venue expert mark the floor ('dance here, avoid the column') and then handing that marked-up plan to the choreographer. That handoff — a spatial annotation, not a full rehearsal — is the core mechanism of MAMHOI. The committed claim: MAMHOI factorizes scene-aware human-object interaction generation through an explicit affordance interface, so that scene understanding and interaction dynamics can be learned from separate, non-paired datasets. This is not the first paper to generate human-object interactions or to reason about scenes, but it is a genuine architectural proposal for decoupling the two problems through a structured intermediate representation — the affordance — rather than requiring expensive paired human-object-scene training data. The data problem this attacks is real and well-known. Human-scene datasets (like PROX, HUMANISE) tell you where people walk and sit in rooms. Human-object datasets (like GRAB, BEHAVE) tell you how hands grip cups or bodies lift boxes. But datasets containing all three — a specific human interacting with a specific object in a specific scene — are vanishingly rare. MAMHOI's factorization means you don't need them. A scene-conditioned model predicts affordances (where and how an interaction can feasibly happen), and a separate affordance-conditioned model generates the actual human-object motion. The affordance representation is the explicit contract between the two modules. Architecturally, this belongs to the family of conditional generative models for human motion — diffusion-based or VAE-based generators conditioned on spatial and semantic inputs. The key structural choice is making the affordance an explicit, interpretable intermediate rather than a latent bottleneck. This is closer to how robotics pipelines factorize perception and planning than to end-to-end approaches that try to learn the full mapping from scene to motion in one shot. The authors evaluate in complex indoor environments and report reduced object-scene penetration (objects passing through walls and furniture) while preserving human-object interaction quality. On the integrity front, the evaluation is simulation-based, with the authors measuring physical plausibility metrics like penetration volume and interaction quality scores. The abstract doesn't name specific numeric baselines or cite head-to-head numbers against named competitors, which limits how confidently we can place this on the ladder. The project page exists but code availability and benchmark standardization are not confirmed from the abstract alone. This is a methods paper with qualitative and quantitative evaluation, not an independent replication or a community challenge result. The milestone question for this line of work is: when does generated scene-aware HOI become good enough for downstream use in embodied AI training, game animation, or VR content creation? The penetration-reduction result matters because penetration artifacts are the single most obvious failure mode that breaks immersion or training signal. But the gap between 'reduced penetration' and 'production-quality motion' is still significant — you'd want sub-centimeter penetration, diverse interaction types, and real-time generation. The obvious experiment not run: testing MAMHOI's factorization in a robotics sim-to-real transfer setting, where a robot uses the affordance predictions to plan real-world object manipulation in cluttered scenes. The authors likely scoped to the vision/graphics community rather than robotics, and the data and evaluation infrastructure for sim-to-real is a different pipeline entirely. This reads as (a) — scope and infrastructure constraints, not a negative result hidden.