Imagine you're teaching someone to cook by showing them three recipes on video. They can't replicate the meals exactly — their kitchen is different, their hands are different — but a good learner would adapt. Now imagine that learner, after each semi-successful attempt, makes tiny variations (different counter, slightly different grip), keeps the ones where the dish actually comes out, and uses those successes as starting points for the next round of experiments. After several rounds, three recipes have become thirty, all edible, all genuinely varied. That's the core mechanism of InterMimicGen. The committed claim: a unified pipeline that takes heterogeneous human motion-capture datasets of object interactions, retargets them to a dexterous humanoid robot in simulation, trains a single generalist tracking policy that can execute all of them, and then closes a self-evolving data flywheel where small task-preserving perturbations compound across rounds into broad behavioral coverage — all without collecting new human demonstrations. The authors consolidate multiple MoCap datasets (GRAB, TACO, OakInk2, InterDance) into a single humanoid reference collection, which alone is a meaningful engineering contribution. The retargeting pipeline is the load-bearing novelty. Prior work either sacrificed hand-object contact fidelity when mapping human motions to robots (different kinematics, different finger counts) or handled locomotion and manipulation separately. InterMimicGen preserves whole-body coordination AND dexterous hand-object relationships simultaneously by using contact-aware optimization that respects the robot's actual joint limits and hand morphology. This is what makes downstream tracking feasible rather than garbage-in-garbage-out. The generalist tracker is a single physics-based policy trained across the full retargeted dataset — covering locomotion, grasping, bimanual manipulation, and their combinations. This contrasts with prior humanoid tracking systems (e.g., PHC, ExBody) that typically specialize in either locomotion OR upper-body manipulation but not dexterous whole-body loco-manipulation at this scale. The policy uses privileged simulation state during training and distills to observation-based control for deployment. The data flywheel is where the self-evolving claim lives. Each augmentation round applies small, task-preserving edits — spatial relocations of where an interaction occurs, kinematic perturbations of how the body moves — fine-tunes the tracker on the candidates, and filters by simulated task completion. Only variants where the object interaction actually succeeds survive to seed the next round. The authors report that executable motion coverage grows with each iteration, and critically, that later rounds produce variants increasingly distant from the originals while maintaining task semantics. This is a form of self-play for manipulation data, not for policy optimization. Integrity-wise, validation is simulation-heavy: Isaac Gym / Isaac Sim with the Fourier GR-1 humanoid as the primary platform. The paper shows retargeting across different robot configurations and reports real-robot transfer on select tasks, but the bulk of quantitative results are simulated. Baselines include ablations (no flywheel, no contact preservation, specialist vs generalist) rather than head-to-head comparisons against named competing systems at equivalent scale — partly because no prior system operates at this exact intersection (whole-body dexterous loco-manipulation from heterogeneous MoCap). The real-robot demos are qualitative proof-of-concept, not systematic benchmarks. The practical bottleneck going forward is whether the flywheel's compounding coverage actually transfers to real-world robustness or just to sim-domain breadth. The gap between 'growing motion library in simulation' and 'reliable deployment in unstructured environments' remains the central unsolved problem for humanoid manipulation. But as a data-generation and policy-training infrastructure, this is a serious piece of systems engineering that moves the field's data scarcity problem from 'we need more human demos' to 'we need better sim-to-real transfer' — a more tractable bottleneck.