Imagine you're an experienced chef trying to teach a gorilla to cook. The gorilla has longer arms, different proportions, and a fundamentally different sense of balance — but it can watch you through a head-mounted camera and, in principle, learn the task. The problem isn't the gorilla's intelligence; it's that your body and its body are mismatched. EgoAlign is the translation layer that makes the gorilla's attempts converge toward your demonstrated behavior, using a physics simulator as a fitting room where arm-length and gait differences get ironed out before anything touches the real kitchen. The committed claim: you can build useful humanoid robot training data entirely from egocentric human demonstrations, without ever teleoperating the robot or collecting physical-robot demonstrations. The framework converts human video into action and state supervision compatible with a continuous whole-body controller by running a three-stage pipeline — scale alignment adjusts for body-proportion mismatches, controller-in-the-loop refinement uses the simulator to check that proposed motions are actually executable, and causal replay reconstructs the missing robot-proprioceptive states that a camera alone can't capture. The result is a dataset suitable for fine-tuning a vision-language-action (VLA) model that deploys zero-shot on a physical humanoid. The ladder here is interesting because the real competition isn't another sim-to-real method — it's teleoperation. The paper's implicit claim is that human demonstration collection is faster and cheaper than expert teleoperation, and they report reduced on-site acquisition time compared to teleop. But the key technical comparison is refinement versus kinematic-only alignment: controller-in-the-loop refinement improves both simulated hand-alignment error and physical pickup success rate over naive kinematic retargeting. The paper does not report comparison against large-scale robot-demonstration approaches like those from Google DeepMind or 1X, which operate in a different data regime entirely. Architecturally, EgoAlign sits at the intersection of motion retargeting, sim-to-real transfer, and VLA fine-tuning. The core insight is that you can use the robot's own dynamics model as a filter: rather than hoping kinematic correspondence is close enough, you literally run the candidate motions through the simulator and let the controller's execution errors tell you where the retargeting is wrong. This is a controller-in-the-loop approach to demonstration adaptation — not new in spirit (sim-to-real has used domain randomization and system identification for years), but the specific application to cross-embodiment human-to-humanoid transfer with egocentric vision is novel. The downstream model is a VLA fine-tuned on the adapted demonstrations. Integrity is mixed. The validation includes real physical robot deployment — a genuine strength over sim-only papers. Tasks include long-range object relocation, navigation to unseen goal positions, and foot interaction, all evaluated on hardware. However, the paper does not report against community benchmarks (there aren't strong standardized ones for humanoid loco-manipulation yet), and success metrics appear to be internally defined. The absence of pre-registration and the small scale of physical experiments leave room for selection effects. No code release is confirmed in the abstract, though a project page exists. The milestone question is sharp for this subfield. Current humanoid manipulation papers typically demonstrate 1-3 tasks in controlled lab settings. The real unlock is sustained, multi-task performance in unstructured environments — think a humanoid reliably executing 20+ household tasks across varied floor plans. EgoAlign demonstrates 3 task categories; the field needs to reach ~20 with >80% success rates before commercial humanoid assistants become plausible. That's probably 3-5 years out if data pipelines like EgoAlign scale. The obvious experiment not run: scaling to a much larger and more diverse set of human demonstrators and environments. EgoAlign's core promise is that human demos are cheap and abundant — but the paper doesn't stress-test this by, say, collecting demos from 50 different people in 20 different rooms. The honest read is (a) and (c): this is a methods paper establishing the pipeline, and the large-scale data experiment is the obvious follow-up that requires more time and compute than a single submission cycle allows. The question of whether the pipeline degrades gracefully with noisier, more diverse human input is left unanswered.